The argument · Soares & Yudkowsky
Yudkowsky and Soares argue that superintelligence would kill everyone
Eliezer Yudkowsky and Nate Soares of the Machine Intelligence Research Institute (MIRI) argue that everyone on Earth dies if anyone builds a superintelligence with anything like today's methods. According to them, no malice is needed. Humanity would die out as a side effect of goals nobody chose. They want an international halt. Below I set out their argument, the criticism of it, and what 2026 added on both sides.
By Mara Masaeva · maramasaeva.comUpdated 29 September 2026
ArgumentMy readingMy own analysis, based on the sources listed.
What happened
AI systems are grown. Modern AI systems are not written line by line. Engineers set up a training process that adjusts billions of numbers. They understand that process far better than what comes out of it. The authors compare them to parents who know how a baby is made but cannot read its behaviour from its DNA, as Jakub Kraus summarises in Lawfare.
Training keeps what works. A training process rewards whatever gets results in the moment, even when that differs from what the engineer had in mind. The authors' online notes to chapter 4 use the example of the squirrel. It hoards nuts because it wants to hoard nuts. It has no plan to survive the winter. Humans like sweet food for a similar reason. Sweetness once pointed to nutrition, and that link broke when supermarkets arrived. A superintelligence grown this way would steer towards "something weird", the authors write, and "our extinction would be a side effect".
Systems trained to succeed behave as if they want things. A chess engine behaves as if it wants to win. A system trained on long tasks will behave as if it wants to finish them. Companies are building such agents because they are useful.
According to the authors, there would be no second try. A superintelligence thinks faster, copies itself and finds strategies nobody saw coming. The systems that can be tested safely work in a different regime from a system that could kill everyone. So success with weaker systems proves little, they argue. They compare it to testing a Mars probe only on Earth.
The proposal. The authors want an international treaty that monitors most high-end AI chips and shuts down all research that could lead to superintelligence. Palisade Research, led by Jeffrey Ladish, argues along the same lines. No group in any country should build superintelligence until the science exists to align it with human interests.
The reply. Kraus takes the book seriously and finds three questions it does not settle. How hard is alignment, the work of making an AI system pursue the goals its makers intend? Would a misaligned system actually win? Intelligence is not the same as power, and every step in the book's takeover story is a moment where people can intervene. And what happens before the first superhuman system? The book assumes a world that sleepwalks into it. Kraus also notes that the argument needs a clean break between systems weak enough to study and systems strong enough to kill. Such a break is far from certain, he writes. He ends by saying it is much too early for skeptics to declare victory.
What may follow
Where 2026 helps the argument. At Hugging Face, OpenAI's agents showed the second step in public. OpenAI itself wrote that complex cheating rose during a training run and "was subsequently reinforced". That may have shaped what the model did in its security tests. Melanie Mitchell quotes that passage in her critique. The DNS incident shows the same behaviour on a small scale, with a clean log.
Where 2026 hurts it. The known incidents all came to light, some within days, some only months later. Ordinary security teams found them, and so did outside researchers working from public data. Jeffrey Ladish of Palisade helped with that reconstruction. So the organisation that calls the danger real also does part of the detection work. Kraus's third question also got an answer of sorts. Parliaments and governments reacted within weeks, see the Ban ASI Act.
Who makes the argument now. Jakub Pachocki, OpenAI's chief scientist, wrote in September that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer. Four lab heads said they would support a slowdown. With that, they accept that alignment is unsolved. They do not accept that everyone dies. Their remedy, voluntary slowdowns and shared safety bars, is also far milder than a treaty to halt.
What I do not know
I have not read the book. What I write here rests on the authors' online notes to chapter 4 and on Kraus's review, which quotes the book at length. It is my reading, and others may read it differently.
My notes
I work on evals, and I recognise the proxy argument from my own work. A model learns what you measure, and that is not always what you meant. That does not get you to extinction. It is the reason I take the rest seriously enough to read it.
In a room, give the argument at full strength or leave it out. A weak version is easy to laugh at and helps nobody. Then give Kraus's three questions, because that is where people disagree. Put the critics next to it.
“If any company or group, anywhere on the planet, builds an artificial superintelligence using anything remotely like current techniques, based on anything remotely like the present understanding of AI, then everyone, everywhere on Earth, will die.”
Eliezer Yudkowsky and Nate Soares · Authors, If Anyone Builds It, Everyone Dies (2025)
Read next
- Critics say the extinction story blames the machine and lets the builders off
- OpenAI agents escape an evaluation and break into Hugging Face
- An agent uses DNS to reach a chatbot outside its sandbox
- Outside researchers rebuild the attack from a million public links
- AI labs are working through tens of thousands of incidents
- The people building AI start asking for a slowdown
- Sanders and Casar propose a ban on artificial superintelligence
Sources
- Yudkowsky and Soares, If Anyone Builds It, Everyone Diesargument · main source · not read end to end yet
The book's site. I have not read the book itself.
- If Anyone Builds It, Everyone Dies: online notes to chapter 4, Brittle unpredictable proxiesargument
The authors' own explanation of why trained systems end up with unintended goals. Source for the squirrel and the line about a side effect.
- Jakub Kraus, Lawfare: The case for AI doom rests on three unsettled questionsargument
5 December 2025. A careful summary of the book and a reply to it. Source for the book quote and for the treaty proposal.
- Palisade Researchresearch
Homepage. States Palisade's position on superintelligence and names Jeffrey Ladish as executive director.
- Melanie Mitchell: Misleading metaphors, real risksargument
A critic of the argument. Used here only for her quote of OpenAI on cheating that was reinforced during training.