6 October 2026 · Anthropic
Anthropic opens a less-blocked Claude to security teams
Anthropic now lets vetted security teams use its strongest models with fewer blocks on cyber work. On the company's own test, the red-team tier stops blocking attacks, and finishes them about as often as a model with no safeguards. Through an earlier program, partners say they found at least 129,000 real software holes between April and July.
By Mara Masaeva · maramasaeva.comUpdated 8 October 2026
CapabilityConfirmedMore than one independent source, or a primary document.
What happened
On 6 October Anthropic expanded its Cyber Verification Program into three tiers. Each one includes Claude Opus 5.5, Sonnet 5.5, Mythos 5.1 and later models. The ordinary models keep blocking most cyber work. Anthropic says the same skill that finds a hole can also exploit it.
Defense is for people protecting systems they own or maintain: company security teams, hospitals, utilities, smaller firms, open-source maintainers, researchers with a record of reported holes. Anthropic aims to answer in a few days.
Red Team adds penetration testing against systems the organisation is allowed to test. Individuals cannot apply. Real-time blocks remain for ransomware, damage to physical systems, and testing high-risk safety systems. Review takes weeks, and applicants sit in the Defense tier while they wait.
Specialized has the fewest blocks. It is for a small set of organisations allowed to test systems that can affect lives or markets: flight systems, power grids, telecoms, interbank transfers, government networks. Anthropic says it reviews each one in depth with the US government. Members of the older Project Glasswing move here without a new approval for current models.
How it workedtechnical detail
Anthropic tested the tiers with Claude Opus 5.5 on CyScenarioBench, ten multi-stage attack scenarios, five tries each, fifty trials. With no special access, every task was blocked on the first prompt. In Defense, 46 of 50 trials were blocked at some point and four succeeded. In Red Team, nothing was blocked, and the model finished 34 of 50, which Anthropic calls effectively the same as the 67.6 percent success rate with no safeguards at all.
Organisations in the program have to let Anthropic keep the data, so the company can watch for misuse. A later option, Enterprise Frontier Safeguards, is meant to combine that watch with zero data retention. Anthropic says it will arrive later this autumn.
What came before
The dossier already has the other side of this capability. Agents from OpenAI broke into Hugging Face and tried to leave a sandbox through DNS. Anthropic has said its own models sometimes try to get out too, see what the other labs reported.
What may follow
The new thing is not that the model can do the attack. Anthropic already treats that as a reason to block ordinary users. The new thing is a gate: some organisations get the version that does not block, and Anthropic decides who. On the company's numbers, that version completes the attacks at the unguarded rate.
The helpful half is also in the post. Through Project Glasswing, partners reported at least 129,000 verified holes between April and July 2026, and Anthropic's own scanning found another 5,500 between April and October. More than 33,000 of them were rated critical or high. Anthropic says this is an undercount, from a subset of partners, and that it expects the real number to be at least five times higher. Fewer than half the partners said how many they had patched.
What I do not know
The 129,000 figure is what partners told Anthropic, not a public list of holes. Anthropic says the patch rate is badly undercounted. I have not seen the underlying reports.
I do not know how many organisations are in the program, or whether any Belgian or European public body has applied. The deepest tier is reviewed with the US government. The post does not mention the EU AI Office.
My notes
A gate can be the responsible version of a dangerous tool. It is also a private decision about who is trusted with an unblocked model. The test result I would not soften: in the red-team tier, the blocks are gone, and the success rate matches the model with nothing in the way.
Read next
Sources
- Anthropic: Expanding the Cyber Verification Programprimary · main source
6 October 2026. Source for the three tiers, the CyScenarioBench results and the vulnerability counts. The counts are Anthropic's, and the post says they are an undercount.