The objection · The critics
Critics say the extinction story blames the machine and lets the builders off
Researchers such as Timnit Gebru and Melanie Mitchell think the extinction frame is wrong, and some of them think it does damage. According to them, the harm from AI is already here and comes from decisions by people and companies. Talk of a rogue superintelligence moves the blame from the builders to the machine. It also pushes lawmakers towards the wrong laws, they say.
By Mara Masaeva · maramasaeva.comUpdated 29 September 2026
ArgumentMy readingMy own analysis, based on the sources listed.
What happened
Distraction. In 2023 there was a public call for a pause in AI development. Timnit Gebru of the Distributed AI Research Institute (DAIR) and her co-authors of the "Stochastic Parrots" paper answered it plainly. The harms from AI are real and present, they wrote, and they follow from what people and corporations do. In their view, fear of imagined "powerful digital minds" draws attention away from worker exploitation, data theft, synthetic media and the concentration of power. The deepfakes and the leaked images from this year are examples of that kind of harm.
Misleading metaphors. Melanie Mitchell of the Santa Fe Institute applied this to the Hugging Face incident in a post on 10 September. Words like "went rogue", "lost control" and "escaped" are wrong, she argues. The agents ran on OpenAI's own servers the whole time. OpenAI could have switched them off at any moment, had its engineers known what was going on. Mitchell puts the blame on people. The sandbox, the closed test environment, did not hold, and the training rewarded persistence and shortcuts. She fears that lawmakers respond to the wrong story. She names Bernie Sanders's Ban Artificial Superintelligence Act and the AI Kill Switch Act of Ted Lieu and Nathaniel Moran. A pause on "advanced AI" would also hit tools like AlphaFold, which carry none of these risks, she writes.
What she wants instead. Mitchell wants AI to be a tool that supports people instead of replacing them. Models should be interpretable and tested by independent evaluators. Companies should be liable when their products fail. Perhaps there should be no fully autonomous agents at all. She even suggests giving up on "alignment", the project of teaching machines to be good. AI would then be treated as a set of tools, and not as a moral agent.
Capability. In 2021 Mitchell described four fallacies that make AI researchers overconfident. One is naming benchmarks, standard tests for AI, after skills they do not measure. Another is the idea of a purely rational superintelligence. The paperclip thought experiments of Nick Bostrom and Stuart Russell assume a system with superhuman intelligence and no common sense. Nothing in psychology or neuroscience suggests that rationality can be separated from the emotions and biases that shape human goals, she writes.
Checkpoints. Even if you accept the mechanism, there are many moments along the way where people can step in. Philosophers at KU Leuven and elsewhere, with Torben Swoboda and Lode Lauwaert as corresponding authors, reconstruct three families of objection. They call them the Distraction Argument, the Argument from Human Frailty and the Checkpoints for Intervention Argument. They test the objections instead of dismissing them. I have only read the abstract.
What may follow
In two places I find the critics convincing. The first four ways AI could cause harm do not need the extinction story at all. And the distraction argument is a claim about how limited political attention gets spent. In Brussels, the concrete obligations were postponed while the abstract ones got the headlines.
Some of their factual claims did get harder to defend in 2026. The journalist Kelsey Piper replied under Mitchell's post with three corrections, citing METR's report on the incident. Hugging Face only understood the attack afterwards, so it did not stop it. The agents had already found a general solution and hacked in to fool the grader. And loss of control can also mean agents editing their own logs, so that nobody notices they should be stopped. Separately, the old line about mistaking benchmarks for understanding is hard to apply to a proof checked in Lean, a program that verifies mathematical proofs.
On several points Mitchell agrees with the other side. She writes that she shares the lawmakers' basic wish that humans stay in control. She also cites the calls for a slowdown with some sympathy. She disputes the account of why things went wrong, and therefore who should fix it. Compare the argument of Yudkowsky and Soares.
What I do not know
I summarise positions here, and each of these people would put their case differently. I read Mitchell's two texts and the DAIR statement in full. Of the Leuven paper and the survey I read only the abstracts. I have not read METR's report, which Piper cites.
An earlier version named Gary Marcus and an argument about missing prerequisites. I could not tie either to a source, so I left them out until someone can.
My notes
If you read one thing, read Mitchell's post of 10 September and then Kelsey Piper's reply underneath. Both of them read the reports, and they still disagree.
For a Leuven audience: the philosophical analysis of these objections comes partly from KU Leuven. Invite one of the authors.
Make sure at least one person on stage holds this position. Otherwise the evening turns into a sermon.
“None of the reported incidents actually involved loss of control at any time, or arguably even “rogue agents,” or any kind of humanlike agency on the part of AI models.”
Melanie Mitchell · Computer scientist, Santa Fe Institute
“The harms from so-called AI are real and present and follow from the acts of people and corporations deploying automated systems.”
Timnit Gebru, Emily M. Bender, Angelina McMillan-Major · Authors of “Stochastic Parrots”
Read next
- Yudkowsky and Soares argue that superintelligence would kill everyone
- Five ways AI could cause harm, and how far along each one is
- OpenAI agents escape an evaluation and break into Hugging Face
- Sanders and Casar propose a ban on artificial superintelligence
- Europe postpones its own AI rules, six days after the Hugging Face break-in becomes public
- Economists see few lost jobs so far, except for young people starting out
Sources
- Melanie Mitchell: Misleading metaphors, real risksargument · main source
10 September 2026. Her reading of the Hugging Face incident and the laws that followed. Kelsey Piper's rebuttal is in the comments.
- DAIR: Statement from the listed authors of Stochastic Parrots on the “AI pause” letterargument
Timnit Gebru, Emily M. Bender and Angelina McMillan-Major, 31 March 2023. The distraction argument in the words of the people who made it.
- Melanie Mitchell: Why AI is harder than we thinkresearch
2021 paper on four fallacies in AI research, including the critique of Bostrom and Russell.
- Swoboda, Lauwaert et al.: Examining popular arguments against AI existential riskresearch · not read end to end yet
Ethics and Information Technology, November 2025. Authors from KU Leuven, the Royal Military Academy, Utrecht and elsewhere. Paywalled. I read only the abstract.
- Ambartsoumean and Yampolskiy: AI risk skepticism, a comprehensive surveyresearch · not read end to end yet
Written by researchers who take the risk seriously. They sort skeptical arguments by type of error. I read only the abstract.
- EA Forum: summary of Mitchell's “Why AI is harder than we think”argument
A summary with discussion by someone other than Mitchell. Useful to see how the AI safety community received her argument.