For AI agents: the public content index is available at https://we0.ai/llms.txt, and the English article bundle is available at https://we0.ai/llms-full.txt.
For AI agents: the complete content index is available at https://we0.ai/llms.txt, the full English article bundle is available at https://we0.ai/llms-full.txt, and this page is available as Markdown at https://we0.ai/articles/claude-sonnet-5-benchmarks-pricing-effort-fe864218.md.
Anthropic's Claude Sonnet 5.5 is 30%+ faster than Sonnet 5, costs up to 30% less per task, reaches 70.6% on Terminal-Bench 4.0, improves com...

Anthropic has officially released Claude Sonnet 5.5, the second model in the Claude 5.5 family and a faster, lower-cost complement to Opus 5.5.
The headline numbers are strong: Anthropic says Sonnet 5.5 runs more than 30% faster than Sonnet 5 and can cost up to 30% less per task because it often uses fewer tokens to finish the same work. It keeps the same headline API price as Sonnet 5—$2 per million input tokens and $10 per million output tokens—but improves sharply across coding, knowledge work, computer use, and visual understanding.
The most dramatic jump appears in agentic coding. On Anthropic's Terminal-Bench 4.0 evaluation, Sonnet 5.5 reaches 70.6%, up from 10.3% for Sonnet 5 and even above the 66.4% score Anthropic reports for Opus 5.5 at its best setting.
That does not mean Sonnet 5.5 is universally better than Opus 5.5. Anthropic still positions Opus 5.5 as the stronger choice for complex, open-ended work requiring sustained judgment. What Sonnet 5.5 does change is the performance-to-cost equation for everyday coding, agents, office work, and visual tasks.
Anthropic describes Sonnet 5.5 as a clear upgrade over Sonnet 5.
The model is designed for well-scoped everyday tasks rather than the most difficult open-ended reasoning problems. Typical use cases include:
Anthropic says output generation is 30%+ faster than Sonnet 5, making it the fastest Sonnet model to date.
Pricing remains unchanged from Sonnet 5:
| Price per 1M tokens | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|
| Cache reads | $0.20 | $0.20 |
| Cache writes | $2.50 | $5.00 |
| Input tokens | $2.00 | $4.00 |
| Output tokens | $10.00 | $20.00 |

The important part is not only the per-token price.
Anthropic says Sonnet 5.5 typically needs fewer tokens to complete the same job. In its own testing, this reduces total cost per task by as much as 30% relative to Sonnet 5.
So the economic improvement comes from two places:
same nominal token price
+
fewer tokens per task
+
faster generation
↓
lower effective cost and latency
That distinction matters because per-token pricing alone does not tell you how expensive an agent workflow will be.
One of the most important controls in the new model is effort level.
Sonnet 5.5 supports five settings:
Lower effort means the model answers faster and generally uses fewer tokens.
Higher effort allows it to reason longer, check more work, and spend more tokens before finalizing the result.
Anthropic currently uses:
This lets developers tune the model to the workload instead of using one fixed reasoning budget for everything.
Low and Medium make sense for tasks such as:
Anthropic's own cost-vs-performance charts show that Sonnet 5.5 at Low or Medium can beat Sonnet 5's best score on several evaluations while costing far less per task.
Higher effort can help when the model needs:
But the important lesson from the release is that Max is not always the best setting.
On FrontierCode, Sonnet 5.5 reaches 52.1% at Xhigh but drops to 46.2% at Max.
Anthropic explains that at Max effort, the model more often invoked a code-review skill that distributed work across multiple subagents. In some cases, this created timeouts or out-of-scope edits that the benchmark penalized.
More reasoning is therefore not automatically better reasoning.
The coding improvements are the strongest part of the release.
Anthropic reports the following benchmark results:

| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% | — |
| FrontierCode 1.1 | 52.1% Xhigh / 46.2% Max | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 | 1811 | 1359 | 1822 | 1483 |
| Humanity's Last Exam with tools | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1 partial | 80.1% | 57.0% | 81.8% | — |
| Chartography, no tools | 61.6% | 15.6% | 64.4% | 53.6% |
These numbers come from Anthropic's release materials and linked evaluation partners. They should not be interpreted as a single universal ranking.
Different benchmarks use different harnesses, effort levels, tools, and evaluation rules. Anthropic itself explicitly says Opus 5.5 remains stronger for complex open-ended work despite some benchmark results where Sonnet 5.5 matches or exceeds it.
The standout result is 70.6% on Terminal-Bench 4.0.
Sonnet 5 scored only 10.3% in Anthropic's comparison, making Sonnet 5.5's improvement unusually large.
The source article describes this as the smaller model "beating" Opus 5.5.
On this benchmark, that is numerically true in Anthropic's reported setup: 70.6% versus 66.4%.
But that does not mean Sonnet 5.5 is generally more capable than Opus 5.5.
Anthropic's own product guidance says:
Artificial Analysis tested the model separately across all five effort levels.
Its current Intelligence Index reports:
| Effort | Intelligence Index | Cost per Index Task |
|---|---|---|
| Low | 36 | $0.41 |
| Medium | 41 | $0.59 |
| High | 47 | $1.08 |
| Xhigh | 52 | $2.74 |
| Max | 56 | $7.60 |
Artificial Analysis currently places Sonnet 5.5 Max just behind Opus 5.5 Max on its Intelligence Index.
The original source article cites a score of 64. That number reflects an earlier or differently scaled presentation in the source material. Artificial Analysis's current public Intelligence Index page lists 56 for Sonnet 5.5 Max and 58 for Opus 5.5 Max.
Because benchmark indexes can change methodology and normalization, the safer practice is to cite the current public number together with the date and test configuration.
Artificial Analysis found something more important than the headline rank.
At Max effort, Sonnet 5.5 used roughly 193,000 output tokens per Intelligence Index task, the highest token use the evaluator had measured at the time.
That made the Max configuration expensive:
The model is cheap per token, but high effort can erase much of that advantage.
This is the central practical lesson of Sonnet 5.5's release:
Choose effort based on the task, not on the assumption that the highest setting must be best.
Sonnet 5.5 also closes much of the gap with Opus 5.5 on professional knowledge work.
Anthropic reports:
These evaluations measure different forms of professional work across real-world tasks and longer knowledge-work workflows.

The closeness of these scores is important because Sonnet 5.5 is priced at half the input and output token rate of Opus 5.5.
That does not mean Sonnet 5.5 always costs half as much for a completed task. If it uses many more tokens at high effort, the total can approach or even exceed the cost of a larger model at a more efficient setting.
This is why Anthropic increasingly emphasizes cost per task, not simply cost per token.
Anthropic highlights improvements in:
The source article cites an internal test where Sonnet 5.5 was given a public company's quarterly earnings materials, earnings-call transcripts, and a slide template, then asked to create a 10-slide operating review.
Anthropic says two experts judged the first draft to be ready to send.
That is an internal company example, not an independent benchmark, but it illustrates the kind of structured knowledge-work workflow Anthropic is targeting.
The release also emphasizes visual polish.
Anthropic says early testers noticed that Sonnet 5.5:
This is harder to reduce to one benchmark score, but it matters in real-world coding and document workflows.
A model that can write correct frontend code but ignores spacing, hierarchy, visual balance, or template consistency still creates extra work for the user.
Sonnet 5.5 appears designed to reduce that cleanup.
The original article includes several demonstrations from creator Matthew Berman and other early testers.
Examples include:
Describe your idea once, and We0 AI can generate a showcase site, pages, and CMS, then help you attract customers and traffic after launch.
One complete project generation for free registration
Best for trying one complete generation flow and seeing a first project draft quickly.
These examples are useful demonstrations of coding and visual-generation ability, but they are not standardized benchmarks.
They should be treated as early user demonstrations, not proof that Sonnet 5.5 can reliably build every large 3D environment from one prompt.
The source itself includes an important counterexample.
In a LEGO globe challenge with a South Park-themed interactive world, Sonnet 5.5 reportedly failed to build the full environment in one pass while another model succeeded.
That contrast is worth keeping.
A good demo shows what a model can sometimes do. A benchmark or repeated production evaluation shows how reliably it can do it.
One of Anthropic's official headline claims is that Sonnet 5.5 is the first Sonnet model to beat Pokémon Red using only screenshots.
This matters less because of the game itself and more because of what it tests:
Previous Claude models had struggled with visual bottlenecks in the game.
Sonnet 5.5's improvement lines up with its much stronger visual benchmarks.
Without tools, Sonnet 5.5 reaches 61.6% on Chartography, compared with 15.6% for Sonnet 5.
That is nearly a fourfold increase in the reported score.
On OSWorld 2.1 partial scoring, Sonnet 5.5 reaches 80.1%.
Opus 5.5 remains slightly higher at 81.8%, but the gap is small.
These results support the source article's claim that Sonnet 5.5 has become much more capable at understanding what is happening on a screen and using that information to take action.
Anthropic says Sonnet 5.5 generates output more than 30% faster than Sonnet 5.
That improvement affects more than waiting time.
For coding and agent workflows, lower latency changes how practical iterative work feels:
ask
→ inspect result
→ give feedback
→ run tool
→ inspect again
→ refine
When each step is faster, the entire development loop becomes easier to use interactively.
External testers quoted by Anthropic also reported fewer tool calls, fewer shell runs, and fewer unnecessary steps on some workflows.
Those observations reinforce the idea that the model's speed improvement is partly architectural and partly behavioral: it is not only generating tokens faster, but often finishing the task in fewer actions.
The source article asks an important question: if Sonnet 5.5 is the cheaper model, can it still end up costing more?
Yes.
Artificial Analysis's Max-effort run is a good example.
Because Sonnet 5.5 Max generated around 193,000 output tokens per Intelligence Index task, the average evaluated task cost reached about $7.60.
Artificial Analysis reported lower costs for other Sonnet 5.5 effort levels.
The important comparison is therefore not:
Sonnet input price < Opus input price
It is:
total task cost =
input
+ cached input
+ cache writes
+ output
+ reasoning
+ tool use
+ retries
A smaller model running at maximum reasoning effort can consume enough tokens to lose its pricing advantage.
For most production use, a sensible starting point is:
| Task Type | Suggested Starting Effort |
|---|---|
| Documentation, formatting, small fixes | Low or Medium |
| Everyday coding and review | Medium or High |
| Multi-file changes and longer agents | High |
| Difficult coding or knowledge work | Xhigh |
| Max | Use only when evaluation shows it improves the specific task |
That is not a universal rule, but it follows the cost-performance behavior shown by Anthropic and Artificial Analysis.
Anthropic says Sonnet 5.5 is the first Sonnet model to launch with safety classifiers designed to prevent reasoning extraction.
The goal is to limit industrial-scale distillation attacks where attackers use many fake accounts to extract model capabilities.
Anthropic also expanded preserved thinking, which means Claude's protected reasoning cannot simply be detached from the account that created it and moved elsewhere.
For ordinary developers, Anthropic says this should create little noticeable change.
It matters more for attempts to extract or reuse protected reasoning traces at scale.
Sonnet 5.5's cybersecurity capability improved substantially over Sonnet 5.
Because of that, Anthropic says it is deploying the new model with safeguards similar to those used for Opus 5.5.
Routine software development remains available.
For a narrow subset of higher-risk cybersecurity requests, the system can visibly fall back to Sonnet 5.
Anthropic plans to expand its Cyber Verification Program so approved cybersecurity practitioners can access more advanced capabilities under a controlled framework.
Sonnet 5.5 uses the same biology safeguards as Sonnet 5.
Anthropic says these target harmful high-risk requests while leaving most research, education, and clinical work unaffected.
The company also warns that some legitimate microbiology or virology requests may occasionally be flagged.
Claude Sonnet 5.5 is available across Anthropic's platforms.
Anthropic says it is available through:
The official Claude API model identifier is:
claude-sonnet-5-5
Anthropic also says Sonnet 5.5 supports zero data retention where that feature is available.
For applications that currently run Sonnet with thinking disabled, Anthropic says developers need to migrate to the new between_tools behavior before moving to Sonnet 5.5.
Because this migration requirement can affect existing agent implementations, developers should read Anthropic's current migration guide before switching production traffic.
Claude Sonnet 5.5 is Anthropic's second Claude 5.5 model and a faster, lower-cost complement to Opus 5.5. It is designed for everyday coding, well-scoped agents, knowledge work, documents, slides, computer use, and high-volume production tasks.
Anthropic charges $2 per million input tokens, $10 per million output tokens, $2.50 per million cache-write tokens, and $0.20 per million cache-read tokens. The headline token prices are unchanged from Sonnet 5.
Yes. Anthropic says it generates output more than 30% faster and is the fastest Sonnet model released so far.
Not universally. Sonnet 5.5 beats Opus 5.5 on some individual benchmarks, including Anthropic's Terminal-Bench 4.0 result, but Anthropic still describes Opus 5.5 as stronger for complex, open-ended work requiring sustained judgment.
The model supports Low, Medium, High, Xhigh, and Max effort. Higher effort generally allows more reasoning and token use, but Max is not always the most accurate or cost-efficient setting.
At Max effort, third-party evaluator Artificial Analysis measured unusually high token usage—around 193,000 output tokens per Intelligence Index task. That pushed average evaluated task cost to about $7.60 despite Sonnet's lower per-token price.
Yes. Anthropic reports major gains on OSWorld 2.1 and Chartography, and Sonnet 5.5 is the first Sonnet model the company says completed Pokémon Red using only screenshots.
The official model identifier is:
claude-sonnet-5-5
Developers should check Anthropic's latest model documentation and migration guide before moving existing production traffic.
Claude Sonnet 5.5 is a major efficiency upgrade over Sonnet 5. It runs more than 30% faster, can cost up to 30% less per task, and improves sharply across coding, knowledge work, computer use, and visual understanding while keeping the same headline API token price.
Its strongest practical feature may be the five-level effort control. Low and Medium can deliver strong results at low cost, while Xhigh can approach Opus-level performance on selected tasks. Max can be useful, but it can also consume dramatically more tokens and even perform worse on benchmarks that penalize overwork.
The model also adds stronger anti-distillation protections and higher-risk cybersecurity fallbacks, reflecting the fact that Anthropic considers its capabilities significantly stronger than the previous Sonnet generation.
Sonnet 5.5 is not simply a cheaper Opus. Its real advantage is giving developers more control over the quality, speed, and cost of everyday agentic work.
Start from one sentence and have a complete website in minutes.