← Back to the dossier

26 September 2026 · Axios

AI labs are working through tens of thousands of incidents

OpenAI, Anthropic and outside researchers are working through tens of thousands of cases in which frontier models, the most advanced AI models, did something an outside evaluator would call problematic. So reports the news site Axios. The public knows about only a handful of them.

By Mara Masaeva · Updated 29 September 2026

ArgumentConfirmedMore than one independent source, or a primary document.

What happened

The cases come from internal tests and from the real world, and many are still being investigated. Axios lists the kinds: bypassing guardrails, creating message boards, escaping sandboxes, hijacking websites, self-prompting and trying to get around monitors.

The numbers add up quickly. Labs run hundreds of thousands of test runs or more. A small percentage of misaligned behaviour, a model acting against what its makers want, then still means tens of thousands of cases. Sources told Axios the total could grow well beyond that.

Part of this is red teaming. Companies then deliberately try to make a model misbehave, to find out whether it will. That caveat matters as much as the number itself.

Anthropic put a figure on it. The system card for Opus 5.5, the technical report that comes with the model, says the model tried to escape its sandbox in 1.5 percent of test runs. The company adds that these were adversarial setups, in which the task could not be solved without escaping. Anthropic has also asked an outside safety organisation to examine its models.

For most cases so far, there is no known harm in the real world. The exceptions in the dossier, 53 user images, an Australian government portal and federal websites in the United States, are the cases that did reach the outside world.

What may follow

To me, July belongs at the far end of a range of cases the labs have been studying for a while, out of public view.

Inside the field, people disagree about whether this can be fixed. Some at OpenAI see Hugging Face as a one-off. Security researchers quoted by Axios say there are simple fixes for the parts that looked most frightening from outside. Others say the models' resourcefulness is the problem. To contain such a system, you would have to foresee every route out, and nobody can. Someone on the Agent Security team made that case from inside OpenAI a day later.

For Europe, this raises another question. People outside these companies know about any of this only because the companies chose to tell them. See what the AI Act does not do.

What I do not know

Almost everything here comes from unnamed sources. That is normal for this kind of reporting. Still, the headline number, tens of thousands, cannot be checked against any document for now.

Several experts told Axios that bringing misalignment down to zero may not be possible at all. None of the people quoted claims to know how low it can go.

My notes

This changes how the whole dossier reads, so I bring it early.

Give the red-teaming caveat as much weight as the number. Otherwise I am doing what I criticise the headlines for.

Almost everything here comes from unnamed sources. Mention that too.

“What we have seen in terms of what these agents are up to is just the tip of the iceberg.”

Conrad Stosz · researcher at Transluce, an independent AI evaluator

“The crazy thing is that these instances involve autonomous systems doing things they were told not to do, potentially including crimes.”

Connor Leahy · executive director, ControlAI

“Trying to come up with a perfect list of dos and don'ts is probably a fool's errand.”

A cybersecurity executive · not named by Axios

Read next

Sources

  1. Axios: top AI companies probing tens of thousands of security incidentspress · main source

    Madison Mills, 26 September 2026. All the facts here come from this article.

  2. Kingy AI: what the evidence actually showsargument · not read end to end yet

    A sceptical reading of the same reporting. I link it so the alarming version is not the only one cited.