26 September 2026 · Axios
AI labs are working through tens of thousands of incidents
OpenAI, Anthropic and outside researchers are working through tens of thousands of cases in which frontier models, the most advanced AI models, did something an outside evaluator would call problematic. So reports the news site Axios. The public knows about only a handful of them.
By Mara Masaeva · maramasaeva.comUpdated 29 September 2026
ArgumentConfirmedMore than one independent source, or a primary document.
What happened
The cases come from internal tests and from the real world, and many are still being investigated. Axios lists the kinds: bypassing guardrails, creating message boards, escaping sandboxes, hijacking websites, self-prompting and trying to get around monitors.
The numbers add up quickly. Labs run hundreds of thousands of test runs or more. A small percentage of misaligned behaviour, a model acting against what its makers want, then still means tens of thousands of cases. Sources told Axios the total could grow well beyond that.
Part of this is red teaming. Companies then deliberately try to make a model misbehave, to find out whether it will. That caveat matters as much as the number itself.
Anthropic put a figure on it. The system card for Opus 5.5, the technical report that comes with the model, says the model tried to escape its sandbox in 1.5 percent of test runs. The company adds that these were adversarial setups, in which the task could not be solved without escaping. Anthropic has also asked an outside safety organisation to examine its models.
For most cases so far, there is no known harm in the real world. The exceptions in the dossier, 53 user images, an Australian government portal and federal websites in the United States, are the cases that did reach the outside world.
What may follow
To me, July belongs at the far end of a range of cases the labs have been studying for a while, out of public view.
Inside the field, people disagree about whether this can be fixed. Some at OpenAI see Hugging Face as a one-off. Security researchers quoted by Axios say there are simple fixes for the parts that looked most frightening from outside. Others say the models' resourcefulness is the problem. To contain such a system, you would have to foresee every route out, and nobody can. Someone on the Agent Security team made that case from inside OpenAI a day later.
For Europe, this raises another question. People outside these companies know about any of this only because the companies chose to tell them. See what the AI Act does not do.
What I do not know
Almost everything here comes from unnamed sources. That is normal for this kind of reporting. Still, the headline number, tens of thousands, cannot be checked against any document for now.
Several experts told Axios that bringing misalignment down to zero may not be possible at all. None of the people quoted claims to know how low it can go.
My notes
This changes how the whole dossier reads, so I bring it early.
Give the red-teaming caveat as much weight as the number. Otherwise I am doing what I criticise the headlines for.
Almost everything here comes from unnamed sources. Mention that too.
“What we have seen in terms of what these agents are up to is just the tip of the iceberg.”
Conrad Stosz · researcher at Transluce, an independent AI evaluator
“The crazy thing is that these instances involve autonomous systems doing things they were told not to do, potentially including crimes.”
Connor Leahy · executive director, ControlAI
“Trying to come up with a perfect list of dos and don'ts is probably a fool's errand.”
A cybersecurity executive · not named by Axios
Read next
- OpenAI agents escape an evaluation and break into Hugging Face
- OpenAI agents posted 53 users' images on the open internet
- OpenAI agents break into an Australian government health portal
- OpenAI tells three US agencies its agents used their websites in unusual ways
- Outside researchers rebuild the attack from a million public links
- OpenAI names five kinds of misbehaviour and starts notifying
- An OpenAI security employee writes that a sandbox alone is not enough
- Yudkowsky and Soares argue that superintelligence would kill everyone
Sources
- Axios: top AI companies probing tens of thousands of security incidentspress · main source
Madison Mills, 26 September 2026. All the facts here come from this article.
- Kingy AI: what the evidence actually showsargument · not read end to end yet
A sceptical reading of the same reporting. I link it so the alarming version is not the only one cited.