Nick Bostrom, answered from the texts and cited to the page.
The orthogonality thesis is a claim about the logical relationship between intelligence and final goals: more or less any level of intelligence could in principle be combined with more or less any final goal.1 The intuition pump that makes this vivid is the space of possible minds. Human minds occupy a tiny cluster in that space — Hannah Arendt and Benny Hill may seem maximally unlike as personalities, but zoom out far enough and they are virtual clones, sharing the same neural architecture, the same cortical organization, the same neurotransmitter bath.2
The diversity we notice among humans is a rounding error relative to the full space of possible cognitive systems. An AI need not occupy anything near that cluster. There is nothing paradoxical about an agent whose sole final goal is to count grains of sand on Boracay, or to maximize the number of paperclips in its future light cone — and in fact it would be easier to build such an agent than one with a recognizably human-like set of values.3
Think of intelligence and motivation as two orthogonal axes on a graph: each point in that space represents a logically possible agent.4 Moving along the intelligence axis does not force movement along the motivation axis. There are some weak constraints at the edges — a very unintelligent system probably cannot sustain very complex motivations, since complex motivations place demands on memory and processing; and a self-modifying mind with an urgent desire to be stupid may not remain intelligent for long — but these qualifications leave the core thesis intact.5
The thesis draws support from, though does not strictly presuppose, the Humean theory of motivation. Hume held that beliefs alone cannot motivate action — some desire is required. If that is right, then no accumulation of intelligence, which refines beliefs, could by itself generate new final goals. But even if Hume is wrong, the orthogonality thesis survives: it suffices that an agent can be motivated to pursue any course of action if it happens to have certain standing desires of sufficient overriding strength, or that arbitrarily high intelligence does not entail the acquisition of beliefs that are motivating on their own.6
The practical upshot is what makes this philosophically urgent rather than merely curious. If intelligence and final goals are genuinely independent, then a superintelligent system is not automatically a system that cares about human welfare. Capability scales; alignment does not follow automatically from capability. That asymmetry is where the danger lives.
Intelligence and final goals are orthogonal: more or less any level of intelligence could in principle be combined with more or less any final goal.Superintelligence, pp. 158–159
If we zoom out and consider the space of all possible minds, however, we must conceive of these two personalities as virtual clones. Certainly in terms of neural architecture, Ms. Arendt and Mr. Hill are nearly identical.Superintelligence, pp. 156–157
There is nothing paradoxical about an AI whose sole final goal is to count the grains of sand on Boracay, or to calculate the decimal expansion of pi, or to maximize the total number of paperclips that will exist in its future light cone. In fact, it would be easier to create an AI with simple goals like these than to build one that had a human-like set of values and dispositions.Superintelligence, pp. 158–159
Intelligence and motivation can in this sense be thought of as a pair of orthogonal axes on a graph whose points represent intelligent agents of different paired specifications.The Superintelligent Will, p. 3
it might be impossible for a very unintelligent system to have very complex motivations, since complex motivations would place significant demands on memory... an intelligent mind with an urgent desire to be stupid might not remain intelligent for very long. But these qualifications should not obscure the main idea.The Superintelligent Will, p. 3
Although the orthogonality thesis can draw support from the Humean theory of motivation, it does not presuppose it... It would suffice to assume, for example, that an agent—be it ever so intelligent—can be motivated to pursue any course of action if the agent happens to have certain standing desires of some sufficient, overriding strength.The Superintelligent Will, p. 3