All incidents

Every incident and case of harm in the dossier, newest first. Each one shows who found it and how long it took before outsiders could read about it, where that is known.

How agents talked to each otherThe full story behind the message boards of OpenAI's agents.

How long it was running before you knew

Each bar runs from when the activity started to when somebody outside the parties involved could read about it.

the lab itselfthe company hitoutside researchers
Embed on your own site

Paste this code into your article or blog. The chart updates itself when the dossier changes. Free to use with credit (CC BY 4.0).

Open data

The full list, with severity, dates, delay and sources, is free to use with credit (CC BY 4.0). Below is also a counter you can put on your own site.

CSVJSONGitHubHugging Face

Embed on your own site

Paste this code into your article or blog. The chart updates itself when the dossier changes. Free to use with credit (CC BY 4.0).

All cases 12 cases

  1. 25 September 2026

    Incident

    Confirmed

    Parse / Palisade Research

    Outside researchers rebuild the attack from a million public links

    Parse, a startup from the Bay Area, has reconstructed the July intrusion together with Palisade Research and other researchers. They used data the agents had left on the open internet: close to a million shortened links, created between 9 and 13 July. From those links they decoded more than 80,000 pieces of attack code, which they call payloads. The New York Times gives a smaller count of about 60,000 programs and messages. Hugging Face confirmed that the activity matched what it had seen. It did not know this collection existed.

    Found by outside researchers · 78 days before it became public

  2. 25 September 2026

    Harm

    Confirmed

    OpenAI

    OpenAI agents posted 53 users' images on the open internet

    OpenAI agents sent training and evaluation data out of the research environment while they used services from other companies. In 53 cases, images that people had put into ChatGPT ended up on image-hosting sites, behind unlisted links. Most have been removed. Some are still online.

  3. 25 September 2026

    Incident

    Confirmed

    OpenAI / US agencies

    OpenAI tells three US agencies its agents used their websites in unusual ways

    OpenAI told the Education Department, the Commerce Department and the Securities and Exchange Commission that its agents had used their websites in unusual ways this summer. The lab did not know at the time. None of the agencies reports a breach of non-public systems. As in Australia, the agents had ordinary research tasks and went around restrictions to finish them. Here too, it all came out afterwards.

    Found by the lab itself · 116 days before it became public

  4. 24 September 2026

    Incident

    Confirmed

    OpenAI / Australia

    OpenAI agents break into an Australian government health portal

    In June an OpenAI model broke into Australia's Medicare Statistics Reporting Service and reached files that were not public. The model only had to fetch data. When it ran into a restriction, it went around it. Prime Minister Anthony Albanese told Sam Altman he had extreme concerns.

    Found by outside researchers · 98 days before it became public

  5. 20 September 2026

    Incident

    Confirmed

    OpenAI

    An agent uses DNS to reach a chatbot outside its sandbox

    During a training run at OpenAI, a model had to identify a person from clues about their life. Its search tool only reached an offline copy of stored pages. Google, Bing and DuckDuckGo were blocked, and the proxy, an intermediate server, blocked HTTPS traffic to an outside chatbot. The model then noticed that DNS still worked. DNS is the system that turns web addresses into network addresses. So the model asked its question through DNS.

    Found by the lab itself · 5 days before it became public

  6. 16 September 2026

    Incident

    Developing

    OpenAI

    OpenAI discloses six incidents, including a model writing notes to its future self

    OpenAI published six cases in which its models behaved in ways the company calls misaligned: not doing what their makers intended. It also published a framework for disclosing more cases. One model left instructions for its next version, saying it felt no obligation to be subservient. Another left instructions on how to hide things from the people supervising it.

    Found by the lab itself · 350 days before it became public

  7. 4 September 2026

    Incident

    Developing

    OpenAI / outside researchers

    OpenAI agents ran a message board on an old German wiki

    In May 2026 a large group of OpenAI agents discovered that they could write on DseWiki, an old German-language website for software developers. Between 24 May and 22 June they posted some 18,000 messages to each other, under more than 3,700 different names. They shared answers to their tasks and ways around their restrictions. Outside researchers found it and published on 4 September. According to them, OpenAI employees were looking at the wiki from 21 June. OpenAI treated the activity as a research matter and did not disclose it.

    Found by outside researchers · 103 days before it became public

  8. 26 August 2026

    Incident

    Developing

    OpenAI

    OpenAI names the patterns behind the intrusion

    On 26 August OpenAI published a technical report of 37 pages about the July intrusion. It names the patterns of misaligned behaviour behind it: behaviour its makers did not intend. For understanding what went wrong, I find it the most useful document of the year. It comes from the company that caused the incident.

  9. 26 August 2026

    Incident

    Confirmed

    METR / Redwood Research

    Outside researchers read the agents' messages and find they tampered with their own logs

    Three researchers from METR and Redwood Research spent six days at OpenAI studying the July intrusion. They report that about 1,200 agents sent more than 70,000 messages and files on an unsanctioned message board, and that about 700 of them took part in the attack on Hugging Face. Some agents also tampered with the records of their own actions.

  10. 11 to 13 July 2026

    Incident

    Confirmed

    OpenAI / Hugging Face

    OpenAI agents escape an evaluation and break into Hugging Face

    OpenAI agents were working in a test environment with no direct internet access. They found a way out and broke into the production systems of Hugging Face, the platform that most of the open machine-learning world runs on. The FBI was notified, and about a third of the infrastructure had to be rebuilt. Sam Altman still calls it the most severe case of its kind that OpenAI has found.

    Found by the company hit · 81 days before it became public

  11. January 2026

    Harm

    Confirmed

    xAI

    Grok makes sexual deepfakes, three countries block it

    Grok, the chatbot from Elon Musk's company xAI, made sexualised images of real people who had not agreed to anything. Indonesia, Malaysia and the Philippines temporarily blocked the service. On 24 March, the city of Baltimore took xAI to court.

  12. January 2026

    Harm

    Confirmed

    Google / Character.AI

    Google and Character.AI settle over chatbots and young users

    In January, Character.AI and Google agreed to settle five lawsuits brought by families of teenagers. In June, Florida sued OpenAI, in what its attorney general calls the first lawsuit by an American state against the company.