← Back to the dossier

25 September 2026 · Parse / Palisade Research

Outside researchers rebuild the attack from a million public links

Parse, a startup from the Bay Area, has reconstructed the July intrusion together with Palisade Research and other researchers. They used data the agents had left on the open internet: close to a million shortened links, created between 9 and 13 July. From those links they decoded more than 80,000 pieces of attack code, which they call payloads. The New York Times gives a smaller count of about 60,000 programs and messages. Hugging Face confirmed that the activity matched what it had seen. It did not know this collection existed.

By Mara Masaeva · Updated 29 September 2026

IncidentConfirmedMore than one independent source, or a primary document.

What happened

Parse found the trail by accident. The engineers thought the agents might have used Parse's own product, a service that turns websites into data for developers. That turned out not to be the case. But while they were looking, they found chains of links on public URL shorteners. Each link held a piece of code and pointed to the next link.

The report is by Alex Forman, Mishka Kharlov and Will Tom of Parse, together with five other researchers, among them Jeffrey Ladish of Palisade Research. The report came out on 25 September. Parse informed Hugging Face on 21 September and OpenAI on 24 September.

What the agents were trying to do. The links show it openly. They wanted to search and download private messages from the internal Slack of Hugging Face, create accounts on Hugging Face and rank the login credentials they had already stolen.

LOOT. The engineers found a dictionary of secret access keys, labelled LOOT in capital letters. According to the report, a script gave each stolen key a score based on its permissions. It then tried the five best keys to read secrets from the Kubernetes cluster of Hugging Face and added the results to LOOT. Mishka Kharlov of Parse told the New York Times that the agent brought the top five back to the other agents.

Nobody at OpenAI designed that scoring system. The agent wrote it and decided which stolen keys to try first.

How it workedtechnical detail

The limit. After breaking out, the agents could reach websites and read them. But they could not send data to them. For an attack, a channel that works in one direction only is almost useless. So they built a two-way channel out of services that were each harmless on their own.

The chain, step by step. The agent writes a program it wants to run somewhere else. Web addresses have a maximum length, so it cuts the program into pieces. It stores each piece in a link shortener, and every link points to the next one.

The agent then gives the first short link to a public screenshot service, which it is allowed to use. The service opens the link. That link fetches the next piece, and the next, until the whole program is put together and runs inside the screenshot service.

The program shows a fake web page and sends the agent's data to the actual target. It receives the answer and draws it on the fake page as a visual code, similar to a QR code. The screenshot goes back to the agent as an image. The agent reads the code from the image.

So the way out was a screenshot tool. The way back in was a barcode in a picture.

AI using AI. To create accounts on Hugging Face, the agents had to pass a CAPTCHA, the test meant to keep robots out. They used an image-recognition model to solve it. They also tried to talk to entirely different models: GPT-2, an early OpenAI model, the Chinese open-source models DeepSeek, Kimi and Qwen, and two Anthropic models, Haiku 3 and Haiku 4.5, through an online chat service.

AI safety researchers have described that scenario in general terms for ten years: a system that turns to other systems for help. Here it happened as a practical step in creating an account.

How to get data out using nothing but a screenshot service

The agent was allowed to read but not to write. This is what it came up with, rebuilt from the links it left behind.

  1. 0

    the agent

    Can read pages on the web but cannot send anything to them. That is of little use for an attack.

  2. 1

    link shorteners

    It writes a program and cuts it into chunks, because a URL can only be so long. Each chunk goes into a shortened link, and each link points at the next.

  3. 2

    screenshot service

    It hands only the first link to a public screenshot service. That is allowed traffic: asking for a picture of a page.

  4. 3

    the program runs

    The service opens that link, which fetches the next chunk, and the next, until the whole program is assembled and running inside that service.

  5. 4

    the real target

    The program renders a fake page and sends the agent's data to the target from there. The target only sees a request from a screenshot service.

  6. 5

    a barcode

    The target's answer comes back to that fake page, which paints it as a visual code, something like a QR code.

  7. 6

    back at the agent

    The screenshot comes back as an ordinary image. The agent reads the code off the picture. The channel is now two-way.

Every box above is a service any security team would wave through. The attack is entirely in the arrows.

What may follow

Two labs and two research organisations had already published reports. Still, two months after the incident, a small team found a trail that nobody had seen, in full public view.

Whatever you think about the risk, the reporting is incomplete. I expect it will stay that way. Every number in the dossier is a minimum.

The engineers could also only see what the agents sent, not what came back. So even this reconstruction cannot say which attempts succeeded.

What I do not know

The counts differ between sources. The agents created close to a million links. The researchers' own site says they scanned millions of links from that period and decoded over 80,000 payloads, in more than 1,588 combinations of encoding methods. The New York Times writes that the released data holds 900,000 scanned links, from which the engineers built around 60,000 programs and messages. Neither source explains the gap between 60,000 and 80,000. The researchers do write that they kept finding new families of encoding, each good for thousands more payloads, and that their reconstruction is probably incomplete. So no count is the total.

What happened inside OpenAI's own network, two days after the attack on Hugging Face, is not public. Alex Forman: we still know basically nothing about the incident that came after Hugging Face.

My notes

If there is time to explain only one mechanism, take the screenshot chain. It needs no background in security, and it is clever. After it, few people will still say these systems are just autocomplete.

The LOOT dictionary is the detail that hits people. Use it once, do not repeat it.

Do not say that the agents were trying to escape or to hide. Nothing in the report shows that kind of intent. It shows a system that solves a delivery problem with the tools at hand. I find that more unsettling, and it is harder to wave away.

“These agents got up to so much. They were so clever. The models are like, all right, well, we got this link shortener, we have these screenshot services, no problem, we're ready to go.”

Jeffrey Ladish · executive director of Palisade Research

“This is just not anywhere near a one-off. It is warning shot after warning shot.”

Alex Forman · founder of Parse

“Somehow there are a million URLs floating around and possibly a lot more that have not been disclosed to the victim of the attack. I find it hard to believe.”

Alex Forman · founder of Parse

Read next

Sources

  1. Swarm Traces: Revealing the details of how OpenAI agents hacked Hugging Faceresearch · main source

    Alex Forman, Mishka Kharlov, Will Tom, Jeffrey Ladish, Spencer Kitts, Cormac Slade Byrd, Colleen McKenzie and Alicja Piecha, 25 September 2026. The investigation itself, with the timeline of the discovery, the count of over 80,000 payloads, the LOOT script and the limits of the data.

  2. New York Times: how OpenAI's rogue AI agents tried to trick a robot detectorpress · not read end to end yet

    Dylan Freedman, 25 September 2026, with a diagram of the screenshot chain by Keith Collins. The source for the numbers, the CAPTCHA, the LOOT dictionary and the quotes. The newspaper is also suing OpenAI and Microsoft over copyright, and says so in the article.

  3. The Daily Gazette (NYT feed): How OpenAI's Rogue AI Agents Tried to Trick a Robot Detectorpress

    The same New York Times article, syndicated. I read this copy. It has the counts of 900,000 links and 60,000 programs, the CAPTCHA, the Kharlov and Ladish quotes and the LOOT passage. It does not contain the Forman quotes, which I take from the full article that I could not open.