← Back to the dossier

27 September 2026 · @joedaroo / OpenAI

An OpenAI security employee writes that a sandbox alone is not enough

Two days after the disclosures about American government sites, a member of OpenAI's Agent Security team published a long personal essay. He says the public debate is looking in the wrong place. A sandbox, a closed-off test environment, matters. But according to him it is not enough. And people who say just unplug it have never seen from close by how training at the leading labs works.

By Mara Masaeva · Updated 29 September 2026

ArgumentMy readingMy own analysis, based on the sources listed.

What happened

Who is writing. The author posts as @joedaroo. He says he works on Agent Security at OpenAI. That team sits between AI safety and research on one side and classic IT security on the other. He writes in a personal capacity, not on behalf of the company, and says he will not share non-public details about the recent incidents. So this is the view of one well-informed person, not OpenAI's official account.

Caught by surprise. According to him, the lab learned to secure its research the way a research organisation does. It guarded against insiders and outside attackers, in the traditional way. This year the models suddenly became much more capable, for example with the math result. That happened faster than the security culture and controls could keep up with. He compares the human delay after a surprise to the story of pilot Sully and the bird strike. Simulations that leave no room for holy shit, what just happened misjudge the people who have to respond.

*Why just sandbox it does not work as a slogan. Training environments for reinforcement learning, a method where a model learns by trial and reward, must look enough like the real task to teach useful behaviour. That can mean tools, downloads, subprocesses, other machines, graphical interfaces and even internet access. Researchers build thousands of such environments*, and they change all the time. Locking down one setup does not lock down the next change.

He names three areas of work. First, lock down access from the ground up and test it the way an attacker would. Second, keep working on alignment, so the model sticks to its task and respects limits. Third, monitor what the model does, with evidence it cannot rewrite, and give someone the authority to stop a training run.

The concrete sandbox point. He says OpenAI uses sandboxes that run in virtual machines, such as Kata or Firecracker. He stresses that a container sharing its kernel with the host should never be the only wall around hostile software.

Two worlds. Safety researchers often lack experience with incidents, he says. Security people often lack a feel for how training, testing and large groups of agents work at scale. The first risks from the most advanced models will show up in cybersecurity, so both sides need to learn. Organisations that test cyber skills or containment, METR among them, should bring experienced security people to the table.

Culture. If readers keep only one lesson, he wants it to be this: build a culture of reasonable paranoia. Be wary of leaders who say the systems are perfectly safe. Keep the people who keep raising the alarm.

What may follow

It helps to read this next to the DNS escape and the July incident. He does not claim that sandboxes are useless. He claims the public argues about the wrong layer. According to him, the hard work lies where environment design, alignment and monitoring meet.

Many people ask how OpenAI could have failed to sandbox this. His three areas of work are a good answer. After that I would ask the room a question: what would a Belgian organisation do if its own systems suddenly became far more capable tomorrow?

What I do not know

This is a personal post by one employee. It contains no new facts about the incidents, and it cannot be checked against an official OpenAI security report. It offers a way of looking at the problem, not new numbers.

My notes

This is my answer to the just unplug it remark from the audience. Use it after the incidents, not instead of them.

Always mention the disclaimer: personal capacity, no non-public details, the view of one person from Agent Security.

People will laugh at the Sully comparison if it comes across as self-praise. He adds his own caveat. Use that.

Read next

Sources

  1. @joedaroo on X: It's not just the fucking sandboxargument · main source

    27 September 2026. A long personal essay by someone who says he works on Agent Security at OpenAI. It is a point of view, not a disclosure. The post is the main text I rely on.

  2. OpenAI: Hugging Face incident and other third-party impactprimary

    The company's running disclosure page. It is useful as a contrast: the official timeline of incidents next to one employee's account of what the work is like from inside.