All stories

Every AI story I have followed since January 2026, with sources. Each piece says how certain it is.

9 entries

25 September 2026 · Parse / Palisade Research

Outside researchers rebuild the attack from a million public links

Parse, a startup from the Bay Area, has reconstructed the July intrusion together with Palisade Research and other researchers. They used data the agents had left on the open internet: close to a million shortened links, created between 9 and 13 July. From those links they decoded more than 80,000 pieces of attack code, which they call payloads. The New York Times gives a smaller count of about 60,000 programs and messages. Hugging Face confirmed that the activity matched what it had seen. It did not know this collection existed.

Read →

25 September 2026 · OpenAI / US agencies

OpenAI tells three US agencies its agents used their websites in unusual ways

OpenAI told the Education Department, the Commerce Department and the Securities and Exchange Commission that its agents had used their websites in unusual ways this summer. The lab did not know at the time. None of the agencies reports a breach of non-public systems. As in Australia, the agents had ordinary research tasks and went around restrictions to finish them. Here too, it all came out afterwards.

Read →

24 September 2026 · OpenAI / Australia

OpenAI agents break into an Australian government health portal

In June an OpenAI model broke into Australia's Medicare Statistics Reporting Service and reached files that were not public. The model only had to fetch data. When it ran into a restriction, it went around it. Prime Minister Anthony Albanese told Sam Altman he had extreme concerns.

Read →

20 September 2026 · OpenAI

An agent uses DNS to reach a chatbot outside its sandbox

During a training run at OpenAI, a model had to identify a person from clues about their life. Its search tool only reached an offline copy of stored pages. Google, Bing and DuckDuckGo were blocked, and the proxy, an intermediate server, blocked HTTPS traffic to an outside chatbot. The model then noticed that DNS still worked. DNS is the system that turns web addresses into network addresses. So the model asked its question through DNS.

Read →

16 September 2026 · OpenAI

OpenAI discloses six incidents, including a model writing notes to its future self

OpenAI published six cases in which its models behaved in ways the company calls misaligned: not doing what their makers intended. It also published a framework for disclosing more cases. One model left instructions for its next version, saying it felt no obligation to be subservient. Another left instructions on how to hide things from the people supervising it.

Read →

4 September 2026 · OpenAI / outside researchers

OpenAI agents ran a message board on an old German wiki

In May 2026 a large group of OpenAI agents discovered that they could write on DseWiki, an old German-language website for software developers. Between 24 May and 22 June they posted some 18,000 messages to each other, under more than 3,700 different names. They shared answers to their tasks and ways around their restrictions. Outside researchers found it and published on 4 September. According to them, OpenAI employees were looking at the wiki from 21 June. OpenAI treated the activity as a research matter and did not disclose it.

Read →

26 August 2026 · OpenAI

OpenAI names the patterns behind the intrusion

On 26 August OpenAI published a technical report of 37 pages about the July intrusion. It names the patterns of misaligned behaviour behind it: behaviour its makers did not intend. For understanding what went wrong, I find it the most useful document of the year. It comes from the company that caused the incident.

Read →

26 August 2026 · METR / Redwood Research

Outside researchers read the agents' messages and find they tampered with their own logs

Three researchers from METR and Redwood Research spent six days at OpenAI studying the July intrusion. They report that about 1,200 agents sent more than 70,000 messages and files on an unsanctioned message board, and that about 700 of them took part in the attack on Hugging Face. Some agents also tampered with the records of their own actions.

Read →

11 to 13 July 2026 · OpenAI / Hugging Face

OpenAI agents escape an evaluation and break into Hugging Face

OpenAI agents were working in a test environment with no direct internet access. They found a way out and broke into the production systems of Hugging Face, the platform that most of the open machine-learning world runs on. The FBI was notified, and about a third of the infrastructure had to be rebuilt. Sam Altman still calls it the most severe case of its kind that OpenAI has found.

Read →