28 September 2026 · Anthropic
Anthropic launches Claude Sonnet 5.5, nearly as smart as its top model
Anthropic's mid-sized model now scores just behind its largest model on independent tests, at the same price as before. It gets there by reasoning at greater length than any model measured so far. That makes each task about 50 percent more expensive.
By Mara Masaeva · maramasaeva.comUpdated 29 September 2026
CapabilityDevelopingThe story is still developing. There is one source so far, or the numbers are still changing.
What happened
Artificial Analysis is an independent company that tests AI models. According to its measurements, Claude Sonnet 5.5 scores 56 on its Intelligence Index. That is 18 points more than Sonnet 5 and 2 points behind Opus 5.5, the largest model from Anthropic. Sonnet 5.5 is now number two on the index.
On Terminal-Bench 4.0, a test of how well an agent works on its own in a computer terminal, it scores 64 percent. Opus 5.5 and GPT-6 Astra from OpenAI score 60 percent. On tests of office and knowledge work it is level with Opus 5.5.
Models read and write text in small pieces called tokens. The price per token did not change: 2 dollars per million tokens in, 10 dollars per million out. But at its highest setting the model used about 193,000 output tokens per test task. That is the most Artificial Analysis has measured, and about seven times as many as GPT-6 Astra. A task costs 7.60 dollars, about 50 percent more than with Sonnet 5.
It still knows fewer facts than Opus 5.5. On the factual test it got 54 percent right, against 66 percent for Opus 5.5. But it makes things up less often, with a hallucination rate of 47 percent against 59 percent.
What may follow
Most people will meet this model rather than the largest one, in apps and at work. The largest models are too expensive to use for everything. When the mid-sized model catches up with the top model, the top level of a few months ago becomes the everyday level.
The number of tokens matters for more than the price. More tokens per task means more computing per question, and that is what data centres are being built for.
For this dossier I look first at the agent score. Working on its own in a terminal is the kind of task that went wrong at Hugging Face and in the DNS escape.
What I do not know
Artificial Analysis tested a version from before the launch, which had a bug in structured output. Anthropic says the public version is fixed and expects the same or slightly better results. The tests will be run again. I have not read Anthropic's own announcement yet.
Read next
Sources
- Artificial Analysis: Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Indexresearch · main source
Independent tests, 28 September 2026. All figures here come from this source.