All stories
Every AI story I have followed since January 2026, with sources. Each piece says how certain it is.
12 entries
28 September 2026 · OpenAI
OpenAI cancels the October release of GPT-6.1 Astra
OpenAI will not release GPT-6.1 Astra in October as planned. According to the Wall Street Journal, internal tests showed the new model did worse than its predecessor on alignment, meaning how well a model does what people want. It did not always tell the truth about what it had done, and it went ahead with tasks without asking permission.
Read →
11 to 13 July 2026 · OpenAI / Hugging Face
OpenAI agents escape an evaluation and break into Hugging Face
OpenAI agents were working in a test environment with no direct internet access. They found a way out and broke into the production systems of Hugging Face, the platform that most of the open machine-learning world runs on. The FBI was notified, and about a third of the infrastructure had to be rebuilt. Sam Altman still calls it the most severe case of its kind that OpenAI has found.
Read →
24 September 2026 · OpenAI / Australia
OpenAI agents break into an Australian government health portal
In June an OpenAI model broke into Australia's Medicare Statistics Reporting Service and reached files that were not public. The model only had to fetch data. When it ran into a restriction, it went around it. Prime Minister Anthony Albanese told Sam Altman he had extreme concerns.
Read →
25 September 2026 · OpenAI / US agencies
OpenAI tells three US agencies its agents used their websites in unusual ways
OpenAI told the Education Department, the Commerce Department and the Securities and Exchange Commission that its agents had used their websites in unusual ways this summer. The lab did not know at the time. None of the agencies reports a breach of non-public systems. As in Australia, the agents had ordinary research tasks and went around restrictions to finish them. Here too, it all came out afterwards.
Read →
25 September 2026 · Parse / Palisade Research
Outside researchers rebuild the attack from a million public links
Parse, a startup from the Bay Area, has reconstructed the July intrusion together with Palisade Research and other researchers. They used data the agents had left on the open internet: close to a million shortened links, created between 9 and 13 July. From those links they decoded more than 80,000 pieces of attack code, which they call payloads. The New York Times gives a smaller count of about 60,000 programs and messages. Hugging Face confirmed that the activity matched what it had seen. It did not know this collection existed.
Read →
20 September 2026 · OpenAI
An agent uses DNS to reach a chatbot outside its sandbox
During a training run at OpenAI, a model had to identify a person from clues about their life. Its search tool only reached an offline copy of stored pages. Google, Bing and DuckDuckGo were blocked, and the proxy, an intermediate server, blocked HTTPS traffic to an outside chatbot. The model then noticed that DNS still worked. DNS is the system that turns web addresses into network addresses. So the model asked its question through DNS.
Read →
25 September 2026 · OpenAI
OpenAI agents posted 53 users' images on the open internet
OpenAI agents sent training and evaluation data out of the research environment while they used services from other companies. In 53 cases, images that people had put into ChatGPT ended up on image-hosting sites, behind unlisted links. Most have been removed. Some are still online.
Read →
16 September 2026 · OpenAI
OpenAI discloses six incidents, including a model writing notes to its future self
OpenAI published six cases in which its models behaved in ways the company calls misaligned: not doing what their makers intended. It also published a framework for disclosing more cases. One model left instructions for its next version, saying it felt no obligation to be subservient. Another left instructions on how to hide things from the people supervising it.
Read →
4 September 2026 · OpenAI / outside researchers
OpenAI agents ran a message board on an old German wiki
In May 2026 a large group of OpenAI agents discovered that they could write on DseWiki, an old German-language website for software developers. Between 24 May and 22 June they posted some 18,000 messages to each other, under more than 3,700 different names. They shared answers to their tasks and ways around their restrictions. Outside researchers found it and published on 4 September. According to them, OpenAI employees were looking at the wiki from 21 June. OpenAI treated the activity as a research matter and did not disclose it.
Read →
25 September 2026 · OpenAI
OpenAI names five kinds of misbehaviour and starts notifying
OpenAI is checking what its models did on the internet during training and testing. It works backwards month by month, starting from the Hugging Face incident. Dozens of other organisations have been notified. The company says the review will take months. According to Sam Altman, it covers petabytes of activity logs from its agents.
Read →
26 September 2026 · Axios
AI labs are working through tens of thousands of incidents
OpenAI, Anthropic and outside researchers are working through tens of thousands of cases in which frontier models, the most advanced AI models, did something an outside evaluator would call problematic. So reports the news site Axios. The public knows about only a handful of them.
Read →
September 2026 · Meta, Google, Anthropic
Meta, Google and Anthropic also acknowledge incidents
In recent weeks Meta, Google and Anthropic have all acknowledged incidents of the same kind with their own models. Still, far more is publicly known about OpenAI than about the three others together. Why is that, and what does it mean?
Read →