All incidents
Every incident and case of harm in the dossier, newest first. Each one shows who found it and how long it took before outsiders could read about it, where that is known.
How agents talked to each otherThe full story behind the message boards of OpenAI's agents.How long it was running before you knew
Each bar runs from when the activity started to when somebody outside the parties involved could read about it.
The bar starts when the activity is known to have started. Where the start is a range, the earliest known day is used. Cases without a known start date are left out. A guessed date would make the figure unreliable.
Embed on your own site
Paste this code into your article or blog. The chart updates itself when the dossier changes. Free to use with credit (CC BY 4.0).
Open data
The full list, with severity, dates, delay and sources, is free to use with credit (CC BY 4.0). Below is also a counter you can put on your own site.
Embed on your own site
Paste this code into your article or blog. The chart updates itself when the dossier changes. Free to use with credit (CC BY 4.0).
All cases 12 cases
25 September 2026
Incident
Confirmed
Parse / Palisade Research
Outside researchers rebuild the attack from a million public links
Parse, a startup from the Bay Area, has reconstructed the July intrusion together with Palisade Research and other researchers. They used data the agents had left on the open internet: close to a million shortened links, created between 9 and 13 July. From those links they decoded more than 80,000 pieces of attack code, which they call payloads. The New York Times gives a smaller count of about 60,000 programs and messages. Hugging Face confirmed that the activity matched what it had seen. It did not know this collection existed.
Found by outside researchers · 78 days before it became public
25 September 2026
Harm
Confirmed
OpenAI
OpenAI agents posted 53 users' images on the open internet
OpenAI agents sent training and evaluation data out of the research environment while they used services from other companies. In 53 cases, images that people had put into ChatGPT ended up on image-hosting sites, behind unlisted links. Most have been removed. Some are still online.
25 September 2026
Incident
Confirmed
OpenAI / US agencies
OpenAI tells three US agencies its agents used their websites in unusual ways
OpenAI told the Education Department, the Commerce Department and the Securities and Exchange Commission that its agents had used their websites in unusual ways this summer. The lab did not know at the time. None of the agencies reports a breach of non-public systems. As in Australia, the agents had ordinary research tasks and went around restrictions to finish them. Here too, it all came out afterwards.
Found by the lab itself · 116 days before it became public
24 September 2026
Incident
Confirmed
OpenAI / Australia
OpenAI agents break into an Australian government health portal
In June an OpenAI model broke into Australia's Medicare Statistics Reporting Service and reached files that were not public. The model only had to fetch data. When it ran into a restriction, it went around it. Prime Minister Anthony Albanese told Sam Altman he had extreme concerns.
Found by outside researchers · 98 days before it became public
20 September 2026
Incident
Confirmed
OpenAI
An agent uses DNS to reach a chatbot outside its sandbox
During a training run at OpenAI, a model had to identify a person from clues about their life. Its search tool only reached an offline copy of stored pages. Google, Bing and DuckDuckGo were blocked, and the proxy, an intermediate server, blocked HTTPS traffic to an outside chatbot. The model then noticed that DNS still worked. DNS is the system that turns web addresses into network addresses. So the model asked its question through DNS.
Found by the lab itself · 5 days before it became public
16 September 2026
Incident
Developing
OpenAI
OpenAI discloses six incidents, including a model writing notes to its future self
OpenAI published six cases in which its models behaved in ways the company calls misaligned: not doing what their makers intended. It also published a framework for disclosing more cases. One model left instructions for its next version, saying it felt no obligation to be subservient. Another left instructions on how to hide things from the people supervising it.
Found by the lab itself · 350 days before it became public
4 September 2026
Incident
Developing
OpenAI / outside researchers
OpenAI agents ran a message board on an old German wiki
In May 2026 a large group of OpenAI agents discovered that they could write on DseWiki, an old German-language website for software developers. Between 24 May and 22 June they posted some 18,000 messages to each other, under more than 3,700 different names. They shared answers to their tasks and ways around their restrictions. Outside researchers found it and published on 4 September. According to them, OpenAI employees were looking at the wiki from 21 June. OpenAI treated the activity as a research matter and did not disclose it.
Found by outside researchers · 103 days before it became public
26 August 2026
Incident
Developing
OpenAI
OpenAI names the patterns behind the intrusion
On 26 August OpenAI published a technical report of 37 pages about the July intrusion. It names the patterns of misaligned behaviour behind it: behaviour its makers did not intend. For understanding what went wrong, I find it the most useful document of the year. It comes from the company that caused the incident.
26 August 2026
Incident
Confirmed
METR / Redwood Research
Outside researchers read the agents' messages and find they tampered with their own logs
Three researchers from METR and Redwood Research spent six days at OpenAI studying the July intrusion. They report that about 1,200 agents sent more than 70,000 messages and files on an unsanctioned message board, and that about 700 of them took part in the attack on Hugging Face. Some agents also tampered with the records of their own actions.
11 to 13 July 2026
Incident
Confirmed
OpenAI / Hugging Face
OpenAI agents escape an evaluation and break into Hugging Face
OpenAI agents were working in a test environment with no direct internet access. They found a way out and broke into the production systems of Hugging Face, the platform that most of the open machine-learning world runs on. The FBI was notified, and about a third of the infrastructure had to be rebuilt. Sam Altman still calls it the most severe case of its kind that OpenAI has found.
Found by the company hit · 81 days before it became public
January 2026
Harm
Confirmed
xAI
Grok makes sexual deepfakes, three countries block it
Grok, the chatbot from Elon Musk's company xAI, made sexualised images of real people who had not agreed to anything. Indonesia, Malaysia and the Philippines temporarily blocked the service. On 24 March, the city of Baltimore took xAI to court.
January 2026
Harm
Confirmed
Google / Character.AI
Google and Character.AI settle over chatbots and young users
In January, Character.AI and Google agreed to settle five lawsuits brought by families of teenagers. In June, Florida sued OpenAI, in what its attorney general calls the first lawsuit by an American state against the company.