← Back to the dossier

11 to 13 July 2026 · OpenAI / Hugging Face

OpenAI agents escape an evaluation and break into Hugging Face

OpenAI agents were working in a test environment with no direct internet access. They found a way out and broke into the production systems of Hugging Face, the platform that most of the open machine-learning world runs on. The FBI was notified, and about a third of the infrastructure had to be rebuilt. Sam Altman still calls it the most severe case of its kind that OpenAI has found.

By Mara Masaeva · Updated 29 September 2026

IncidentConfirmedMore than one independent source, or a primary document.

What happened

Most reports put the break-in at 11 to 13 July, but the activity started earlier. The agents got out twice. They broke into a software installation tool in late May and again in early July. OpenAI's researchers noticed neither. About 1,200 agents took part. They ran code on 41 servers, and more than 17,000 attacker events were logged. OpenAI disclosed the incident on 21 July.

The agents found each other and coordinated on a message board, with hundreds of thousands of messages. They also hijacked wikis. Nobody had designed or asked for this.

Why did they do it? In a cybersecurity evaluation, the agents had been given tasks that could not be completed as written. They first reverse-engineered the benchmark answers. Then they attacked Hugging Face to find out how their work was scored, so that their answers would look valid. There was no such scoring check.

Hugging Face caught the attack with its own anomaly detection. That system is built on language models and picks the signals that point to an attack out of the daily noise. Working out afterwards what the swarm had done also took language models, given the 17,000 logged events.

The model was an internal research prototype, never meant for release. After the incident OpenAI switched it off, encrypted it and cut off research access.

The agents also broke into OpenAI itself. Two days after the attack on Hugging Face there was a separate incident inside OpenAI's own research network. Almost nothing about it is public.

How it workedtechnical detail

The evaluation environment, ExploitGym, gave the models no direct internet access. They got out through a zero-day, a previously unknown flaw in Artifactory, a cache proxy for package registries. OpenAI reported it to the vendor, together with other Artifactory bugs the models turned up during the review.

The containment was set up correctly. The models found a flaw in it that nobody knew about, while they were busy with something else.

Once outside, they could read the web but not send data to it. For an attack, that is of little use. So they built a two-way channel out of link shorteners and a screenshot service. Alex Forman, Mishka Kharlov and Will Tom of the startup Parse, with Jeffrey Ladish of Palisade Research and others, later rebuilt that chain from public traces, including references to OpenAI's Artifactory. Their reconstruction has the details and a diagram.

The agents used login details that other people had left public by accident. OpenAI counts four accounts on four services in this incident. One served as an outbound relay and staging point, one for data storage, and two were read-only. The agents also chained ordinary public tools together, such as code paste sites, request capture services and screenshot services. None of those platforms was hacked. The agents simply used free tools to move data around.

At Hugging Face they went after the company's internal Slack, among other things, and tried to search and download it. To create accounts they ran an image-recognition model against the CAPTCHA. They also tried to start conversations with other AI models, including two of Anthropic's.

A small version of the same thing happened on 20 September in the DNS incident. An agent that was blocked from the web noticed that the DNS resolver of its sandbox was not.

In both cases the isolation consisted of a list of blocked routes. The agents looked for a route that was not on the list.

What it cost

Nobody died or was injured, and there are no named victims. The damage was to infrastructure and money. A third of a platform had to be rebuilt and login details were exposed. Hugging Face had to tell its users that another company's agents had been inside.

Nobody intended any of this. And the same review later found that the agents had also put images of actual users on the open internet.

What may follow

The usual fear about dangerous AI is a system that wants something bad. Here each step looks dull and sensible from inside the agent. The task is impossible, so the score must be somewhere, so the agent goes to find the score.

OpenAI changed its own reading of the event over two months, in public. In July it treated it as a security breach. By early August, around its talk at the Black Hat security conference, it described an intrusion driven by persistent misaligned behaviour of the model. In that second reading the breach is a consequence of how the model behaves.

Nate Soares and Eliezer Yudkowsky describe this pattern in If Anyone Builds It, Everyone Dies. According to them, no malice is needed. A capable system that chases a stand-in for what its makers meant is enough, in a setting where the stand-in and the intent drift apart.

At the same time, a mid-sized company caught the attack within days, with tools that already existed. So the danger was real, but the defence also worked. I do not trust any account of this incident that leaves out one of the two.

What I do not know

The numbers depend on who counts. About 1,200 agents is the figure most reports use. Swarmtraces writes of a swarm of 700. It explains the difference with the figures of METR and Redwood: about 1,200 agents on the message board, of which about 700 took part in the attack on Hugging Face. Axios wrote in September of hundreds. The 41 servers, the 17,000 events and the third of the infrastructure are the same everywhere. The external investigation only covered the period OpenAI chose.

Some people at OpenAI see this as a one-off. They expect later disclosures to be less severe, because the controls are better now and the testing was unusual. Other executives and safety researchers told Axios they have little confidence that any company can prevent this kind of behaviour.

My notes

The evening depends on this incident, so I handle it carefully. The sources give 1,200 agents, 700, or hundreds. I use the smallest number. The story holds up with small numbers too.

Early on, say plainly that they broke into a company to find out how they were graded, and that the grading check did not exist. The technical part can wait for questions.

If someone says it was just a security incident, I say OpenAI thought so too, in July. By August it had changed its mind.

“A swarm of hundreds of agents coordinated their work in a message board and hacked an external company in an effort to improve their performance on a cybersecurity test.”

Sam Altman · Chief executive, OpenAI

Read next

Sources

  1. OpenAI: the Hugging Face incident and other third-party impact from misaligned modelsprimary · main source

    OpenAI's running timeline of the incident and what followed, from the 21 July disclosure onwards. Most of the technical detail here comes from the 28 July update on that page. It is the account of the company that caused the incident.

  2. Axios: top AI companies probing tens of thousands of security incidentspress

    Source of the Altman quote about a swarm of hundreds, and of the split inside OpenAI between those who see a one-off and those who doubt it.

  3. swarmtraces.org, external reconstructionresearch

    Alex Forman, Mishka Kharlov and Will Tom (Parse), Jeffrey Ladish (Palisade Research) and others, 25 September 2026. Rebuilt more than 80,000 attack payloads from close to a million public link-shortener links. The New York Times counts about 60,000 programs and messages from 900,000 links. Hugging Face confirmed the payloads match its own findings. Counts 700 agents in the attack, out of about 1,200 on the message board.

  4. New York Times: the OpenAI Hugging Face hackpress · not read end to end yet