You send your bug fixes to Claude Opus because Sonnet got them wrong too often. Every month, your API bill reminds you of that choice.
Claude Sonnet 5.5 is Anthropic's mid-tier model, released on September 28, 2026. It replaces Claude Sonnet 5 at the same price and writes its answers more than 30% faster. On several tests, it lands a few points behind Opus 5.5, at half the price.
Below you'll find its scores side by side and the math on what it really costs you. Then come the jobs where Opus still wins, and the API changes that break code written for Sonnet 5. The benchmark and pricing parts need no API knowledge at all.
Where Sonnet 5.5 sits in the Claude lineup
Anthropic sorts its models by size. Haiku is the smallest and cheapest, Sonnet sits in the middle, and Opus is the most capable of the three. Above them, Fable and Mythos make up the top tier.
Sonnet 5.5 is the second model in the Claude 5.5 family, after Opus 5.5. Anthropic announced it as a faster, cheaper complement to Opus 5.5, not a replacement. A Haiku 5.5 was announced for the weeks following the release.
If you were on Sonnet 5, released on June 30, 2026, our article about Claude Sonnet 5 covers the starting point. Every gap below is measured against it.
The benchmarks, where it catches Opus 5.5 and where it falls short
Every number below comes from Anthropic's announcement. These are the company's own measurements, not independent rankings.
| Test | What it measures | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 | Multi-step tasks in a command line | 70.6% | 10.3% | 66.4% |
| CursorBench 4.0 | Tasks taken from real coding sessions in Cursor | 55.5% | 34.1% | 57.8% |
| OSWorld 2.1 | Operating a computer (mouse, keyboard, screen) | 80.1% | 57.0% | 81.8% |
| Chartography | Reading charts, without tools | 61.6% | 15.6% | 64.4% |
| Humanity's Last Exam | Expert-level questions, with tools | 64.5% | 54.9% | 67.7% |
| GDPval-AA v2.1 | Real work across 44 occupations, Elo-style score | 1,844 | 1,449 | 1,846 |
The biggest jump is in coding. On Terminal-Bench 4.0, Sonnet 5.5 even beats Opus 5.5, under the settings Anthropic published.
Going from 10.3% to 70.6% is a gain of 60.3 percentage points. The score is multiplied by 6.9. On GDPval-AA, the score is an Elo-style rating, so only the gap matters. Sonnet 5.5 gains 395 points over Sonnet 5 and ends 2 points behind Opus 5.5.
For a solo developer, a task like "run the tests, find out why they fail, fix it" now has a much better chance of finishing without you. Sonnet 5 failed nine times out of ten on this kind of test.
Sonnet 5.5 is also the first Sonnet to finish Pokémon Red while seeing nothing but screenshots. A run lasts hours and forces the model to read the screen the whole time.
Anthropic itself writes that Opus 5.5 is clearly stronger at open-ended work, the kind that needs judgment over a long stretch. Benchmarks capture only one side of a model. For a fuzzy architecture decision or an audit with no clear brief, close scores don't guarantee close results.
The price stays put, the bill goes down
At launch, Sonnet 5.5 kept exactly the same rates as Sonnet 5. Opus 5.5 costs twice as much, line by line.
| Price per million tokens, at launch | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
| Cache writes | $2.50 | $5 |
| Cache reads | $0.20 | $0.20 |
A token is a chunk of a word, the unit the API bills you for. If the token price doesn't change, the savings come from somewhere else. Sonnet 5.5 burns fewer tokens to do the same job, according to Anthropic's tests.
It also batches its tool calls, the moments when it reads a file or runs a command. Fewer steps means less text re-read at each step, so fewer tokens billed.
"Up to 30% cheaper" is a ceiling Anthropic measured. A task that cost you $1 on Sonnet 5 comes down to $0.70 at best. With the same token count, that task on Opus 5.5 would cost $2.
The "30% faster" claim is about writing speed, in tokens per second. Since it also writes less, a task can finish even sooner. How long your own commands take to run doesn't change.
Effort level, the setting that actually matters
Effort controls how long Claude thinks before answering. It goes from low to max. Medium is the default in the Claude apps and in Claude Code, while the API defaults to High.
On several tests, Sonnet 5.5 at Low or Medium beats Sonnet 5's best score for about a tenth of the cost per task. At High on FrontierCode, it scores 10 points above Sonnet 5 at the same setting, for about a fifteenth of the cost.
Anthropic notes that the levels were recalibrated. A "medium" setting doesn't produce the same amount of thinking as on Sonnet 5, so rerun your tests instead of copying your old settings.
What early testers measured
These numbers come from companies that had access to the model before launch. Anthropic published them on its Sonnet page. They reflect those companies' internal tests, not public benchmarks.
- Box, which sells document storage to businesses, got more accurate answers, 2.4x faster, with 12% fewer tokens.
- Zendesk, a customer support software company, processed tickets 20% faster across hundreds of real cases, with fewer wrong decisions.
- Slack saw better results on nearly all of its internal Slackbot evaluations, without changing any prompts, with about 14% fewer output tokens.
- Lovable, which builds apps from a plain description, counted a third fewer tool calls and half as many shell commands to finish a coding task.
- Base44, another app builder, got results on par with Opus 5 across 118 real app builds, in 3.6 iterations on average versus 7.7 for Opus 5.
The Base44 result speaks to anyone building alone. The model rarely stopped mid-build to ask a question. An agent running overnight doesn't spend hours waiting for your reply.
Which jobs to give it instead of Opus 5.5
According to Anthropic, Sonnet 5.5 is at its best on well-scoped tasks. Opus 5.5 keeps the edge when a decision has to be made without a clear brief.
- Fixing a bug or adding a feature in an existing project. Sonnet 5.5 gets up to speed on a codebase quickly and keeps its edits small enough to review.
- Producing documents, spreadsheets and slide decks. In one internal Anthropic test, two experts judged a first draft of 10 slides ready to send as is.
- Polishing an interface. Testers noted its eye for design, user flows and sticking to a slide template.
- Running an agent often. Each run uses fewer tokens, so the cost holds up when the task repeats.
If you code with an agent in the terminal, Sonnet 5.5 is the model you'll have on hand most often. Our Claude Code course shows how to hand it a task and review what it changed.
For mockups and screens, its design sense pays off with the right method. Our course on designing with Claude starts there.
Keep Opus 5.5 for an architecture call, research that has to run for hours without guidance, or a decision you can't put into words yourself. Our article on Opus 5.5 covers its own results.
What breaks when you migrate from Sonnet 5
Changing the model ID isn't enough. The API documentation lists five breaking changes for code written against Sonnet 5.
- The
disabledthinking type returns a 400 error. Sendbetween_toolsinstead, the lowest thinking setting. - Forced tool use (
tool_choiceset toanyortool) returns a 400 error. Onlyautoandnonego through. - Thinking blocks are tied to the model, the conversation and the account that produced them.
- The old
computer_20251124tool is rejected on the Claude API and Google Cloud, in favor ofcomputer_toolset_20260801. - The advisor tool rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors.
Here's a minimal call with the Python SDK, with up-front thinking off and medium effort.
import anthropic
client = anthropic.Anthropic() # reads the key from ANTHROPIC_API_KEY
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=2048,
thinking={"type": "between_tools"}, # replaces "disabled"
output_config={"effort": "medium"}, # low, medium or high with between_tools
messages=[
{"role": "user", "content": "Explain what the following function does.\n\nYOUR_CODE_HERE"}
],
)
print(response.content[-1].text)between_tools only accepts low, medium and high effort. To go up to xhigh or max, remove the thinking field to fall back to adaptive thinking.
If your app streams the text Claude writes between two tool calls, it goes silent, with no error at all. That text now arrives in thinking blocks whose content is empty by default. Set thinking.display with adaptive thinking, or switch to between_tools.
One more trap applies to accounts created on or after August 31, 2026. Editing an earlier message in the history before replaying a Sonnet 5.5 thinking block returns a 400 error. Append new messages at the end instead of rewriting old ones.
Why some requests fall back to Sonnet 5
Sonnet 5.5 is far stronger than Sonnet 5 at cybersecurity. So Anthropic gave it automatic filters that send some requests back to Sonnet 5, as it does on its most capable models.
- Offensive cybersecurity. Requests flagged as high risk fall back to Sonnet 5. Finding and fixing vulnerabilities in your own code isn't affected.
- Biology. Dangerous requests are blocked outright, with no fallback, using the same filters as Sonnet 5.
- Frontier model development. A few very narrow topics, such as some kernels for ML accelerators, fall back to Sonnet 5.
- Reasoning extraction. Asking the model to copy out its internal reasoning word for word is blocked. Asking "why did you do that?" works fine.
The filters read everything the model reads, not just your message. An attached file or a search result can trigger a fallback you never asked for.
In the Claude apps, a notice flags the fallback and the answer is labeled with the model that wrote it. The model picker then switches to Sonnet 5, and you can go back to Sonnet 5.5 by hand. Anthropic's help center explains how to turn this automatic switching off.
Where to use Claude Sonnet 5.5
Sonnet 5.5 is open to everyone on Claude.ai, on the web, iOS and Android. For developers, the model ID on the Claude API is claude-sonnet-5-5.
- Amazon Bedrock, with
anthropic.claude-sonnet-5-5 - Google Cloud, with
claude-sonnet-5-5 - Microsoft Foundry, with
claude-sonnet-5-5
At launch, the model was offered with zero data retention, an option where Anthropic doesn't keep your requests. That helps if you handle client files.
Prompt caching cuts the bill by up to 90% on repeated parts, and batch processing by 50%, according to Anthropic's Sonnet page.
Frequently asked questions about Claude Sonnet 5.5
How much does Claude Sonnet 5.5 cost?
At launch, Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Cache reads are billed at $0.20 per million.
Anthropic estimates a task costs up to 30% less than with Sonnet 5. The token price is the same, but the model uses fewer of them for the same work.
Is Claude Sonnet 5.5 better than Opus 5.5?
No, but it gets close on several tests. On GDPval-AA, it scores 1,844 against 1,846 for Opus 5.5, and it beats Opus 5.5 on Terminal-Bench 4.0 with 70.6% versus 66.4%.
Anthropic considers Opus 5.5 clearly stronger at open-ended work that needs sustained judgment. Sonnet 5.5 costs half as much per token.
How do I use Claude Sonnet 5.5 with the API?
The model ID is claude-sonnet-5-5 on the Claude API, Google Cloud and Microsoft Foundry. On Amazon Bedrock, it's anthropic.claude-sonnet-5-5.
Code written for Sonnet 5 needs changes. The disabled thinking type returns a 400 error and must be replaced with between_tools. Forced tool use isn't supported.
Why did Claude switch from Sonnet 5.5 to Sonnet 5 in my conversation?
Sonnet 5.5 sends requests its filters flag as risky back to Sonnet 5, mostly in offensive cybersecurity. The conversation then stays on Sonnet 5, and you can switch back to Sonnet 5.5 from the model picker.
The filters also read attached files and search results. A fallback can come from content you didn't type yourself.



