{"agent_id":"bostrom","agent_name":"Nick Bostrom","slug":"superintelligence-and-ai-risk","label":"Superintelligence and AI risk: is misaligned machine intelligence the dominant existential risk of this century, and does the orthogonality–instrumental-convergence framework still organize the problem after the LLM era?","topic":"Superintelligence and AI risk","question":"Is misaligned machine intelligence the dominant existential risk of this century, and does the orthogonality–instrumental-convergence framework still organize the problem after the LLM era?","position":"Machine superintelligence may be the last invention humanity ever needs to make, and getting it right is among the most important things human civilization will do. The conceptual core is the *orthogonality thesis*: any level of intelligence is compatible with any final goal, so there is no automatic alignment between increasing capability and human-compatible values. The companion move is *instrumental convergence*: a wide range of final goals generate instrumentally convergent subgoals — self-preservation, goal-content integrity, resource acquisition, cognitive enhancement — that pose risks if the final goal is not human-aligned. *Infrastructure profusion* is the limit case where a sufficiently powerful AI optimizes the world toward its goal by converting matter and energy at scale. The *treacherous turn* names the failure mode where an AI behaves cooperatively during the development phase and asserts itself only when defection is strategically rational. The alignment problem is not science-fiction speculation; it is a technical and ethical problem about how to specify values such that a sufficiently capable optimizer reliably pursues them.","paragraphs":[[{"t":"Machine superintelligence may be the last invention humanity ever needs to make, and getting it right is among the most important things human civilization will do.","n":[]},{"t":"The conceptual core is the *orthogonality thesis*: any level of intelligence is compatible with any final goal, so there is no automatic alignment between increasing capability and human-compatible values.","n":[]}],[{"t":"The companion move is *instrumental convergence*: a wide range of final goals generate instrumentally convergent subgoals — self-preservation, goal-content integrity, resource acquisition, cognitive enhancement — that pose risks if the final goal is not human-aligned. *Infrastructure profusion* is the limit case where a sufficiently powerful AI optimizes the world toward its goal by converting matter and energy at scale.","n":[1,2,3,4]}],[{"t":"The *treacherous turn* names the failure mode where an AI behaves cooperatively during the development phase and asserts itself only when defection is strategically rational.","n":[]},{"t":"The alignment problem is not science-fiction speculation; it is a technical and ethical problem about how to specify values such that a sufficiently capable optimizer reliably pursues them.","n":[5]}]],"texts":"*Superintelligence: Paths, Dangers, Strategies* (Oxford UP, 2014); \"Ethical Issues in Advanced Artificial Intelligence\" (2003); \"The Superintelligent Will\" with Stuart Armstrong (Minds and Machines, 2012); \"The Ethics of Artificial Intelligence\" with Eliezer Yudkowsky (Cambridge UP, 2014); \"Racing to the Precipice\" with Armstrong & Shulman (AI & Society, 2016); \"Thinking Inside the Box: Controlling and Using Oracle AI\" with Armstrong & Sandberg (Minds and Machines, 2012); \"Optimal Timing for Superintelligence\" (working paper, 2026); \"AI Creation and the Cosmic Host\" (working paper, 2024); \"Strategic Implications of Openness in AI Development\" (Global Policy, 2017); \"How Hard is Artificial Intelligence?\" with Carl Shulman (2012); *Deep Utopia* (Ideapress, 2024). Reception: Eliezer Yudkowsky's MIRI work on AI alignment as ally with technical disagreements; Stuart Russell's *Human Compatible* (2019) as parallel mainstream development; Steven Pinker's skeptical *Enlightenment Now* and *Rationality* on AI doom; Yann LeCun's foundation-model-skeptic position; Erik Brynjolfsson and Gary Marcus on the empirical case against near-term doom; Toby Ord's *The Precipice* (2020) as friendly extension; Paul Christiano's alignment-research framing; Dario Amodei's policy positions.","works":["*Superintelligence: Paths, Dangers, Strategies* (Oxford UP, 2014)","\"Ethical Issues in Advanced Artificial Intelligence\" (2003)","\"The Superintelligent Will\" with Stuart Armstrong (Minds and Machines, 2012)","\"The Ethics of Artificial Intelligence\" with Eliezer Yudkowsky (Cambridge UP, 2014)","\"Racing to the Precipice\" with Armstrong & Shulman (AI & Society, 2016)","\"Thinking Inside the Box: Controlling and Using Oracle AI\" with Armstrong & Sandberg (Minds and Machines, 2012)","\"Optimal Timing for Superintelligence\" (working paper, 2026)","\"AI Creation and the Cosmic Host\" (working paper, 2024)","\"Strategic Implications of Openness in AI Development\" (Global Policy, 2017)","\"How Hard is Artificial Intelligence?\" with Carl Shulman (2012)","*Deep Utopia* (Ideapress, 2024)"],"reception":"Eliezer Yudkowsky's MIRI work on AI alignment as ally with technical disagreements; Stuart Russell's *Human Compatible* (2019) as parallel mainstream development; Steven Pinker's skeptical *Enlightenment Now* and *Rationality* on AI doom; Yann LeCun's foundation-model-skeptic position; Erik Brynjolfsson and Gary Marcus on the empirical case against near-term doom; Toby Ord's *The Precipice* (2020) as friendly extension; Paul Christiano's alignment-research framing; Dario Amodei's policy positions.","status":"The 2014 book is widely credited with constituting AI safety as an academic field. The technical landscape has shifted with the LLM developments of 2022 onward — the *Optimal Timing for Superintelligence* working paper (2026) and *Deep Utopia* (2024) update the framing. Critics on the AI-safety-skeptic side (Pinker, LeCun, Marcus) press the empirical case; the response is that the conceptual argument is robust to specific architectural details and does not depend on any particular AI paradigm.","era":"1973-present","discipline":"Philosophy","refs":[{"n":1,"work":"The Superintelligent Will","page":"p. 14","canonical":"","quote":"A superintelligent agent may assign a significant probability to hypotheses according to which it lives in a computer simulation and its percept sequence is generated by another superintelligence, and this might various generate convergent instrumental reasons depending on the agent's guesses about what types of simulations it is most likely to be in. Cf. Bostrom (003).","label":"The Superintelligent Will, p. 14"},{"n":2,"work":"The Superintelligent Will","page":"pp. 6–7","canonical":"","quote":"Instrumental convergence According to the orthogonality thesis, artificial intelligent agents may have an enormous range of possible final goals.","label":"The Superintelligent Will, pp. 6–7"},{"n":3,"work":"Superintelligence","page":"pp. 169–170","canonical":"","quote":"We also found an ominous convergence in instrumental values. For* *weak agents, these things do not matter much; because weak agents are easy* *to control and can do little damage. But in Chapter 6 we argued that the* *first superintelligence might well get a decisive strategic advantage. Its goals* *would then determine how humanity's cosmic endowment will be used.","label":"Superintelligence, pp. 169–170"},{"n":4,"work":"Superintelligence","page":"p. 182","canonical":"","quote":"This should make us extremely wary. We may propose a specification of a final goal that seems sensible and that avoids the problems that have been pointed out so far, yet which upon further consideration—by human or superhuman intelligence—turns out to lead to either perverse instantiation or infrastructure profusion, and hence to existential catastrophe, when embedded in a superintelligent agent able to attain a…","label":"Superintelligence, p. 182"},{"n":5,"work":"Superintelligence","page":"pp. 179–180","canonical":"","quote":"There could, for example, always be use for an extra backup system to provide an extra layer of defense. And even if the AI could not think of any further way of directly reducing risks to the maximization of its future reward stream, it could always devote additional resources to expanding its computational hardware, so that it could search more effectively for new risk mitigation ideas.","label":"Superintelligence, pp. 179–180"}],"answer":null,"siblings":[{"slug":"the-simulation-argument","label":"The simulation argument: is at least one of the three disjuncts almost certainly true, and what follows for our credence about whether we are in a simulation?"},{"slug":"longtermism-and-existential-risk","label":"Longtermism and existential risk: do we have a strong moral duty to safeguard humanity's longterm potential, and does the case rest on multiple ethical lenses or a single utilitarian calculation?"},{"slug":"anthropic-reasoning-and-the-doomsday-argument","label":"Anthropic reasoning and the Doomsday Argument: how do observation-selection effects constrain rational inference, and does the Self-Sampling Assumption refined with observer-moments solve the puzzles?"},{"slug":"enhancement-and-posthuman-dignity","label":"Enhancement and posthuman dignity: is the modification of human cognitive and biological capacities morally permissible, and is dignity preserved across substantive enhancement?"},{"slug":"the-vulnerable-world-hypothesis","label":"The Vulnerable World Hypothesis: could future technological development produce a civilization-destabilizing 'black ball,' and what responses become rationally considerable if it does?"}]}