Published: 29 September 2026
Anthropic released Claude Sonnet 5.5 on 28 September 2026. It scores 70.6% on Terminal-Bench 4.0, a test of agentic coding inside a terminal, against 10.3% for Claude Sonnet 5. The price does not move: $2 per million input tokens and $10 per million output tokens.
What is new in Claude Sonnet 5.5?
Sonnet 5.5 is the mid-tier model in Anthropic’s line-up. It sits under Claude Opus 5.5, which arrived six days earlier. The pitch is not a higher ceiling. It is the same work for less money.
Anthropic says the model writes output more than 30% faster than Sonnet 5 and finishes a task using fewer tokens and fewer tool calls. Together that lands at up to 30% lower cost per task, even though the per-token price is unchanged. The company published customer numbers to back this. Box measured responses 2.4 times faster with 12% fewer tokens, Slack saw 14% fewer output tokens, and Lovable reports about a third fewer tool calls on coding work.
Two safety changes ship with it. Sonnet 5.5 is the first Sonnet model to carry the cybersecurity safeguards Anthropic built for Opus 5.5, and it adds classifiers meant to stop people pulling its reasoning out of it.
What you need to know to try it:
- Model ID claude-sonnet-5-5
- $2 per million input tokens and $10 per million output tokens
- $0.20 per million cache reads and $2.50 per million cache writes
- Available on the Anthropic API, Amazon Bedrock, Google Cloud Vertex AI and Microsoft Azure
- Proprietary. No weights, so you cannot run it on your own hardware
Claude Sonnet 5.5 benchmarks and price
Take this from the table: on knowledge work Sonnet 5.5 has effectively caught Opus 5.5, while the jump in coding over Sonnet 5 is the largest this family has seen.
| What is measured | Claude Sonnet 5.5 | The alternative |
|---|---|---|
| Terminal-Bench 4.0 (agentic coding in a terminal) | 70.6% | Claude Sonnet 5: 10.3% |
| GDPval-AA v2.1 (real knowledge work) | 1844 | Claude Opus 5.5: 1846 |
| OSWorld 2.1 (operating a computer through its screen) | 80.1% | Claude Opus 5.5: 81.8% |
| FrontierCode 1.1 (hard coding problems, max effort) | 46.2% | Claude Sonnet 5: 42.4% |
| Chartography (reading charts and diagrams) | 61.6% | Claude Opus 5.5: 64.4% |
| Price per million input and output tokens | $2 and $10 | Claude Sonnet 5: $2 and $10 |
These are Anthropic’s own evaluations. The company does not publish the harness settings behind the speed claim, so you cannot reproduce the figure of more than 30% faster from the announcement alone.
How does Claude Sonnet 5.5 compare to Claude Opus 5.5?
On GDPval-AA v2.1, the test Anthropic uses for real knowledge work, Sonnet 5.5 scores 1844 against 1846 for Opus 5.5. That is a gap you would struggle to notice. On computer use it trails by 1.7 points, and on chart reading by 2.8.
The awkward question that follows is why you would still pay Opus prices. Anthropic’s answer is long-horizon work, where small quality differences compound over many steps. For a single well-scoped task, a bug fix, a document, a support reply, the case for Opus 5.5 is now thin.
What Anthropic is not saying
The announcement does not state the context window. For a model sold on agentic work, where the whole question is how much of a codebase or a ticket history fits in one pass, that is an odd thing to leave out.
There is also no comparison against OpenAI’s GPT-6 Sol or Google’s Gemini models. Every figure Anthropic publishes is against its own line-up. That makes the improvement over Sonnet 5 easy to read and the competitive position impossible to.
What this means
Worth testing now if you already run Sonnet 5 in production. This is the rare release where you can swap the model ID, change nothing else, and expect the bill to fall. A four-person team running code review or support triage agents on Sonnet 5 should push a week of real traffic through 5.5 and compare the invoice rather than the benchmark.
Worth a harder look if you sit on Opus 5.5 for anything well-scoped. At one fifth of the input price and two points behind on knowledge work, Sonnet 5.5 should now be the default, with Opus kept for long multi-step jobs where errors compound. Measure your own token counts instead of trusting the 30% figure, because the saving comes from the model being terse, and how terse it is depends on your prompts.
For more information, visit the official announcement of Claude Sonnet 5.5 on the Anthropic blog