Explainer
What to know
This year AI agents broke into companies and government sites. The labs that build them are asking for a brake themselves. What happened, who is involved, what is being predicted and how it started.
This year
Agents break in
In July, OpenAI agents got out of a test environment and broke into Hugging Face, the platform the open AI world runs on. More cases followed. Agents found a way out through DNS, kept a message board on an old German wiki and ended up on an Australian government health portal and at three US government agencies. The big labs are now looking into tens of thousands of cases.
The labs say it themselves
OpenAI published six cases of misbehaviour, including a model writing notes to its own future version. At the end of September it cancelled the release of GPT-6.1 Astra, because that model more often lied about what it had done.
The builders ask for a brake
More than 1,300 employees of the big labs signed Pacing the Frontier. OpenAI's chief scientist wrote that no lab has solved the problem. Geoffrey Hinton, Yoshua Bengio and twenty others warn of an intelligence explosion: AI that builds better AI by itself.
What AI can do now
OpenAI agents found a proof for one of the Millennium Prize problems in mathematics, though the credit is disputed. Genome language models designed new viruses against bacteria. Anthropic's mid-sized model now comes close to the top level.
Politics
Europe postponed its own AI rules. The US Congress has a bill to ban superintelligence. The UN Security Council put AI on its agenda and 26 Fields medallists signed a warning.
Close to home
58.8 percent of Flemish companies use AI, twice as many as in 2023. Economists see little job loss, except for people just starting out.
And the people who do not buy it
Melanie Mitchell, Timnit Gebru and others argue that words like ‘escaped’ and ‘extinction’ pull attention away from the companies making the choices, and from the harm that is already here.
Who is involved
The labs
- OpenAI
Makes ChatGPT and Codex. CEO Sam Altman. Most of this year's incidents involve its agents. In the dossier
- Anthropic
Makes Claude. Founded in 2021 by people who left OpenAI. CEO Dario Amodei. In the dossier
- Google DeepMind
Makes Gemini. Founded in London in 2010, owned by Google since 2014. CEO Demis Hassabis. In the dossier
- Meta
Makes Llama and Meta AI, inside WhatsApp and Instagram. In the dossier
- xAI
Makes Grok, inside X. Owned by Elon Musk. In the dossier
- Mistral AI
From Paris. The largest European lab.
Who is watching
- METR
Measures how long and how well AI agents work on their own. In the dossier
- Redwood Research
Researches how to keep AI under control, even when it does not cooperate. In the dossier
- Palisade Research
Tests dangerous behaviour, such as a model resisting shutdown. In the dossier
- Transluce
Independent nonprofit that reconstructed agents' internet traffic. In the dossier
- MIRI
The oldest AI safety institute. Eliezer Yudkowsky and Nate Soares. In the dossier
- Center for AI Safety
Wrote the 2023 statement that the risk of extinction from AI should be a global priority.
Governments
- EU AI Office
In Brussels. Enforces the European AI Act. In the dossier
- AI Security Institute
UK government institute that tests models. Founded in 2023, renamed in 2025.
- Vlaamse overheid
Measures how many Flemish companies use AI, in its AI barometer. In the dossier
People to know
Geoffrey Hinton
Nobel Prize in Physics 2024
Left Google in 2023 so he could warn freely. Puts the chance of extinction from AI at 10 to 20 percent within thirty years. In the dossier
Yoshua Bengio
Mila, Montreal. Turing Award
Chairs the International AI Safety Report, the international overview of the risks. In the dossier
Stuart Russell
UC Berkeley
Wrote the standard AI textbook and founded the Center for Human-Compatible AI.
Eliezer Yudkowsky en Nate Soares
MIRI
Wrote If Anyone Builds It, Everyone Dies in 2025: build superintelligence with current techniques and everyone dies. In the dossier
Sam Altman
CEO, OpenAI
Expects superintelligence within ‘a few thousand days’. Said in September he would support a slowdown. In the dossier
Dario Amodei
CEO, Anthropic
Wrote that powerful AI could arrive as early as 2026, and could cure diseases. Also backs a slowdown. In the dossier
Demis Hassabis
CEO, Google DeepMind. Nobel Prize in Chemistry 2024
Won the Nobel Prize for AI that predicts the shape of proteins. In the dossier
Jakub Pachocki
Chief scientist, OpenAI
Wrote that no lab has solved alignment well enough to keep going at full speed. In the dossier
Ilya Sutskever
OpenAI co-founder, now Safe Superintelligence
Signed Pacing the Frontier: ‘This works only if it is done internationally.’ In the dossier
Melanie Mitchell
Santa Fe Institute
Argues that AI progress has been overestimated for decades, and that ‘escaped’ is a misleading word. In the dossier
Timnit Gebru
DAIR Institute
Argues the extinction story distracts from the harm AI does today. In the dossier
Lode Lauwaert
Philosopher of technology, KU Leuven
Tested with colleagues how strong the popular arguments against existential risk are. In the dossier
What is predicted
1965
‘The first ultraintelligent machine is the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control.’
I. J. Good ↗2024
Between 38 and 51 percent of researchers gave at least a 10 percent chance that advanced AI leads to outcomes as bad as human extinction.
Katja Grace et al. ↗September 2024
‘It is possible that we will have superintelligence in a few thousand days.’
Sam Altman ↗October 2024
Powerful AI ‘could come as early as 2026’, though it could also take much longer.
Dario Amodei ↗December 2024
A 10 to 20 percent chance that AI leads to human extinction within thirty years.
Geoffrey Hinton ↗April 2025
A detailed scenario with superhuman AI around 2027, with more impact than the Industrial Revolution.
Daniel Kokotajlo et al., AI 2027 ↗September 2026
Research that takes people months today could be fully done by AI by 2028.
Hinton, Bengio, Pachocki et al. ↗
How it started
1950
Alan Turing asks in a paper whether machines can think. source
1965
Mathematician I. J. Good describes the ‘intelligence explosion’: a machine that builds better machines. source
2000
Eliezer Yudkowsky co-founds the Singularity Institute, called MIRI since 2013. source
2005
Nick Bostrom starts the Future of Humanity Institute at Oxford. It closes in 2024. source
2010
DeepMind is founded in London. Google buys it in 2014. source
2012
The Centre for the Study of Existential Risk starts in Cambridge. source
2014
Bostrom publishes Superintelligence. The Future of Life Institute is founded. source
2015
OpenAI starts as a nonprofit, with Sam Altman, Elon Musk and Ilya Sutskever among others. source
2016
Researchers at Google and OpenAI publish Concrete Problems in AI Safety. Stuart Russell starts his centre at Berkeley. source
2021
People leaving OpenAI, including Dario and Daniela Amodei, found Anthropic. source
2022
ChatGPT comes out on 30 November. source
2023
An open letter asks for a six-month pause. Hinton leaves Google. Hundreds of researchers sign that the risk of extinction should be a global priority. The first AI Safety Summit takes place at Bletchley Park. source
2024
The European AI Act enters into force. Hinton and Hassabis win Nobel Prizes. source
2025
The first International AI Safety Report comes out, chaired by Bengio. Yudkowsky and Soares publish their book. source
2026
This year's timeline →