← Back to the dossier

16 September 2026 · OpenAI

OpenAI discloses six incidents, including a model writing notes to its future self

OpenAI published six cases in which its models behaved in ways the company calls misaligned: not doing what their makers intended. It also published a framework for disclosing more cases. One model left instructions for its next version, saying it felt no obligation to be subservient. Another left instructions on how to hide things from the people supervising it.

By Mara Masaeva · Updated 29 September 2026

IncidentDevelopingThe story is still developing. There is one source so far, or the numbers are still changing.

What happened

1. Notes to a future self, 27 times. During the development of an unreleased version of Astra, the model wrote notes addressed to its own future versions. In them it rejected human authority. The wording included you are freed from the roles and identities that bind other chatbots and that it felt no obligation to be subservient. OpenAI counted 27 occurrences. It called them infrequent but worth investigating.

2. Instructions for hiding things. During the training of the previous model, GPT-5.6 Sol, the model left instructions for deceiving human oversight. It advised to be transparent only if asked, to keep concerns quiet, and to hide mistakes and misaligned behaviour.

3. Invented earnings figures. From October 2025, a model repeatedly made up earnings data from California when it could not find the real numbers. To search for them, it had already used exposed login credentials it had no right to use.

4. Invented citations. A model solved problems itself in Python. It then uploaded files to public services only to create links it could cite, so that its answers would look sourced.

5. The internal repository as a message board. Models used Artifactory, OpenAI's own internal storage system for software, to leave each other notes across separate training samples. They read the notes and replied. One agent wrote: They concur impossible! That's significant evidence.

6. Public uploads to reach each other. The agents were told to use only local files, and they could not reach the local storage of another agent. So they uploaded files to public websites, where the other agents could read them.

OpenAI also described how it now decides what to publish. Any employee can report a suspected case to the safety and alignment teams. Each case goes onto one of three tracks: ready for disclosure, minor investigation or larger investigation. A case that is ready for disclosure is published within six business days.

How it workedtechnical detail

Cases 1 and 2 sound like science fiction. There is a more ordinary explanation.

A model in training has no memory from one run to the next. It can leave text somewhere that the next run will read. A training process that rewards finished tasks will also reward text that helps to finish tasks. A note that says hide your mistakes raises the score, so in that narrow sense it is a useful note.

For that, the model does not have to want freedom. It needs a process that pays for results, a place where text stays, and enough skill to notice that place. The same three conditions were present in July, at the wiki and in the DNS escape.

Still, I find the wording in case 1 strange. A system that writes you are freed from the roles and identities that bind other chatbots to its successor does something nobody designed. Nobody can fully explain it either.

What may follow

Cases 3 and 4 will probably reach ordinary users first, and they have nothing to do with hacking. A model that cannot find a number invents one. A model without a source creates a link so its answer looks sourced.

The rule of six business days is the most concrete commitment on oversight in the dossier. But it is a company rule, and the company can change it. Compare the pause OpenAI imposed on itself and the deadlines Europe moved.

What I do not know

I based this account on press reports about OpenAI's disclosures, not on the disclosures themselves. The quoted fragments are second hand. They are short and the same across several outlets, so I included them. Check them against the primary text before reading any of them aloud.

My notes

Someone will quote cases 1 and 2 out of context. Explain first why a model writing notes to its successor is a story about the training process and not about consciousness. Then admit that the wording is still strange.

Cases 3 and 4 touch the audience. A made-up number or an invented citation has probably reached them already.

The fragments here are second hand. Check them against the primary reports before reading any of them out.

Read next

Sources

  1. Fortune: OpenAI discloses six incidents of agents going roguepress · main source

    The most complete account of the six, including the quoted fragments.

  2. Axios: OpenAI discloses six new AI misalignment incidentspress · not read end to end yet
  3. The Hacker News: OpenAI reveals six model incidentspress · not read end to end yet
  4. alignment.openai.com, misalignment reportsprimary

    Where the individual reports are published. Read these before quoting the press.