Six days after Anthropic put Claude Opus 5.5 at the top of its lineup, the mid-tier model has closed most of the gap at half the price. Claude Sonnet 5.5 beats Opus 5.5 on one benchmark and trails it by only a couple of points elsewhere, all while keeping Sonnet 5's pricing.
We've seen this pattern before. Claude Sonnet 5 launched, nearing Opus 4.8's agentic benchmarks at a fraction of the cost. Before that, Sonnet 4.6 landed within striking distance of Opus 4.6 on the same terms. Each time, a mid-tier model arrives days or weeks after the flagship and closes most of the gap at half the price.
This is what a logarithmic scaling curve looks like from the outside. Performance rises roughly with the log of compute, so cost rises roughly exponentially with performance. A flagship model sits at the top of that curve, where it's spending heavily to buy the last few points. A cheaper model built weeks later, benefiting from the same underlying advances, only needs to reach a lower point on the curve to look nearly as good. In other words, most of the curve's value is still cheap to get, and only the final stretch is expensive.
I think that's worth parsing out because the leading companies do a bit of wizardry with PR and branding.
TL;DR
- Sonnet 5.5 is Anthropic's new mid-tier model, replacing Sonnet 5 at the same price.
- The headline gain is agentic terminal coding, where it now actually edges past Opus 5.5.
- On knowledge work and computer use, it sits just behind Opus 5.5 for half the per-token price.
- Make it your default Claude model and keep Opus 5.5 for open-ended, judgment-heavy work.
What Is Claude Sonnet 5.5?
Anthropic's model documentation describes Claude Sonnet 5.5 as "the best combination of speed and intelligence" in the current lineup. It is the mid-tier model in Anthropic's Claude 5.5 family, sitting between the flagship Claude Opus 5.5 and the upcoming Claude Haiku 5.5. (We will keep you posted on Haiku 5.5, although Haiku 5 never arrived.)
Compared with Sonnet 5, the jump is larger than a half-version number suggests. I don't think this is surprising after we saw that the jump between Opus 5.5 and Opus 5 was larger than expected.
As for the launch post - Anthropic reports gains on every benchmark it published, with the biggest leaps in terminal coding, chart reading, and knowledge-work evaluations.
Interestingly, Anthropic says this Sonnet model is the first to beat Pokémon Red working only from screenshots. It sounds unusual but this is a test that Anthropic uses as a proxy for long-horizon work and image understanding.
So, Opus 5.5 remains, in the company's words, clearly stronger at complex, open-ended work that needs sustained judgment. But Sonnet 5.5 is the model for tasks where you know what "done" looks like, which is a helpful heuristic.
Claude Sonnet 5.5 Key Features
Most of what's new in Sonnet 5.5 is about doing the same work faster and with fewer tokens, plus a few changes that API users need to plan around.
Get near-Opus results on everyday tasks at Sonnet speed
You can hand Sonnet 5.5 the kind of work that previously wanted Opus, and get it back quicker. Anthropic says output generation is more than 30% faster than Sonnet 5, and that the model uses fewer tokens per task - "up to 30% less per task" - is the figure they use.
Early testers quoted in the launch post back this up with their own numbers:
- Base44 reports 3.6 iterations per app build with Sonnet 5.5, against 7.7 with Opus 5.
- Box saw its workloads run 2.4x faster with 12% fewer total tokens.
- Zendesk measured 20% faster ticket processing.
- Slack reports roughly 14% fewer output tokens.
These are vendor-selected customer quotes. Still, there's a pattern, which is one of fewer iterations and fewer tokens, which matters for your bill, alongside the per-token rate.
Run coding agents in the terminal and on the desktop
Sonnet 5.5 is built to work as an agent: running shell commands, fixing bugs across a repo, and operating a desktop through screenshots. It's live in Claude Code from launch, and Anthropic describes its computer use as close to Opus 5.5.
The practical upshot for many people is that Sonnet 5.5 is now a credible default for Claude Code sessions and for agent loops you build yourself with our Claude API guide.
Produce polished documents, slides, and interfaces
Anthropic singles out a "sharp eye for design" as a Sonnet 5.5 strength. In practice, that should mean better-looking user interfaces when you ask it to refine front-end code, and more finished output when you ask for slides, spreadsheets, or long documents.
The design claim is the one most in need of independent checking. It's the hardest to benchmark, and the launch post supports it with only customer quotes.
Expect safeguards on security and biology work
Sonnet 5.5 is the first Sonnet model to launch with cybersecurity safeguards similar to those on Opus 5.5. When a request is classified as higher-risk security work, the request falls back to Sonnet 5. Biology safeguards are the same as on Sonnet 5, and Anthropic says most software development and life sciences work will be unaffected.
It's also the first Sonnet to ship with classifiers that block reasoning extraction, and thinking blocks are now tied to the account and conversation that produced them. If you run security toolin , test your prompts before switching a pipeline.
Migrate from Sonnet 5 with a few code changes
Swapping the model ID isn't quite enough. Anthropic's documentation lists five breaking changes for code already running on Sonnet 5:
-
Adaptive thinking is on by default, and the lowest setting is
between_tools, which turns off up-front thinking. -
Forced tool use returns an error.
-
Thinking blocks are tied to the model and conversation that produced them.
-
The older
computer_20251124computer use tool is rejected on the Claude API and Google Cloud. -
The advisor tool no longer accepts Opus 4.8, Opus 4.7, or Sonnet 5 as advisors.
There's a quieter change too: text between tool calls now comes back in thinking blocks. An app that streams text to users will go silent between tool calls until you update its display setting. Setting temperature, top_p, or top_k to a non-default value also returns a 400 error. These are specific things you might discover.
How Does Claude Sonnet 5.5 Perform on the Benchmarks?
Claude Sonnet 5.5 beats Sonnet 5 on every benchmark Anthropic published, beats Opus 5.5 on terminal coding, and trails Opus 5.5 by a few points everywhere else. Against OpenAI's GPT-6 Sol, it leads on the knowledge-work and chart-reading evaluations but trails on FrontierCode. All figures below come from Anthropic's launch post.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% | Not reported |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | Not reported |
| FrontierCode 1.1 (Main) | 46.2% (max effort) | 42.4% | 54.4% | 49.3% (52.1% at xhigh) |
| GDPval-AA v2.1 (Elo) | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 (Elo) | 1811 | 1359 | 1822 | 1483 |
| Humanity's Last Exam (with tools) | 64.5% | 54.9% | 67.7% | Not reported |
| OSWorld 2.1 | 80.1% | 57.0% | 81.8% | Not reported |
| Chartography (no tools) | 61.6% | 15.6% | 64.4% | 53.6% |
The table tells a simple story: a big generational step over Sonnet 5, and a near-tie with Opus 5.5. Let's look at what each area means for your work.
Agentic coding
Terminal-Bench 4.0 is where Sonnet 5.5 separates itself. It scores 70.6%, up from Sonnet 5's 10.3% and ahead of Opus 5.5's 66.4%. The benchmark measures whether a model can complete real tasks in a command-line environment, such as multi-step shell workflows, so it maps closely to how Claude Code is used.

A 60-point jump in one release is unusual enough. The rest of the coding picture is less dramatic.
Knowledge work
On GDPval-AA v2.1, which rates model output on economically valuable tasks from real occupations using an Elo-style score, Sonnet 5.5 scores 1844. That's 2 points behind Opus 5.5 and nearly 400 ahead of Sonnet 5.
For teams using Claude for reports, analysis, and business documents, this is the most useful result in the table. On this kind of work, Opus 5.5's extra cost buys only a little.
Reasoning and visual understanding
Chartography, a test of reading and reasoning over charts without tools, shows the biggest relative gain: Sonnet 5.5 scores 61.6% against Sonnet 5's 15.6%. That lines up with Anthropic's claims about improved image understanding.
Computer use
On OSWorld 2.1, which measures whether a model can complete tasks by operating a real desktop operating system, Sonnet 5.5 reaches 80.1%, up from 57.0% for Sonnet 5 and within 2 points of Opus 5.5.
Which Tier Should You Use?
For most developers, Sonnet 5.5 should now be the default Claude model. Opus 5.5 earns its higher price only when the task is open-ended enough that judgment, not speed, is the bottleneck.
The current lineup, per Anthropic's model comparison table, looks like this. All four models except Haiku 4.5 share a 1M-token context window and 128K max output.
| Model | Price (input / output per 1M tokens) | Latency | Best for |
|---|---|---|---|
| Claude Fable 5.1 | $10 / $50 | Slower | Top-end work where capability matters more than cost |
| Claude Opus 5.5 | $4 / $20 | Moderate | Complex, open-ended work needing sustained judgment |
| Claude Sonnet 5.5 | $2 / $10 | Fast | Well-scoped coding, agents, and document work |
| Claude Haiku 4.5 | $1 / $5 | Fastest | High-volume, cost-sensitive tasks |
Within Sonnet 5.5, the effort setting is the other lever. Anthropic recalibrated five levels: low, medium, high, xhigh, and max. Lower settings favor speed and cost, while higher ones let the model reason longer and check its own work. The default is medium in the Claude apps and Claude Code, and high on the API.
One detail from the launch post changes how I'd use it: at low or medium effort, Sonnet 5.5 beats Sonnet 5's best score on several benchmarks for about a tenth of the cost per task. In sum:
-
Bug fixes, refactors, and Claude Code sessions: Sonnet 5.5 at the default effort.
-
High-throughput agent loops and pipelines: Sonnet 5.5 at
lowormedium, then raise effort only on failures. -
Hard coding problems with a clear spec: Sonnet 5.5 at
xhighormaxbefore paying for Opus. -
Ambiguous, multi-day projects: Opus 5.5.
Testing Claude Sonnet 5.5: Hands-On Examples
The benchmarks above are Anthropic's own, so I ran one build task against Claude Sonnet 5.5 and Claude Sonnet 5 with an identical prompt and identical tooling.
The task: turn FiveThirtyEight's recent college grads dataset (173 U.S. majors, CC BY 4.0) into a single-file HTML dashboard that helps a future student compare categories on earnings and employment and find the highest- and lowest-earning majors.
The real test is in the data, which has a few charateristics the prompt never mentions:
- The top earner, Petroleum Engineering, has just 36 survey respondents, and 76 of the 173 majors have fewer than 100
- Earnings are skewed, so a single median bar hides a lot of spread
- One major, Food Science, has blank headcount fields
- The job-type columns don't add up to the number of people employed
Noticing these issues is the point of the test. The test probes Anthropic's claims about having a "sharp eye for design."
The main findings:
- Only Sonnet 5.5 shipped a working dashboard. Sonnet 5 linked to a Chart.js file that doesn't exist on cdnjs (the URL returns a 404), so the library never loaded and every panel, including its plain tables, rendered empty. Sonnet 5.5 also checks whether its chart library loaded and keeps its tables working if it didn't.
- Sonnet 5.5 caught the small-sample trap, mostly. It marks majors with fewer than 30 respondents "treat with caution" and adds a minimum-respondents filter. Its default view puts Petroleum Engineering at #1 without a warning.
- Sonnet 5 would have missed it anyway. Its code reads the sample-size and salary-range columns, then never uses them, and ranks majors by median alone.
- The smaller traps caught nobody out. Both handled the blank Food Science row explicitly, and Sonnet 5.5 computed its college-job share without implying the job columns add up to total employment.
Here's Sonnet 5.5's dashboard with the filter set to majors with at least 100 respondents.

I scored Sonnet 5.5's dashboard 4 out of 5 for honest data handling, 5 for analytical usefulness, and 4 for chart and layout quality. Sonnet 5's blank page I'm including below.
So why did Sonnet 5 fail? The problem was, the the dashboard needed a charting library. Sonnet 5 asked cdnjs for a version it doesn't host. It then got a "404" error. Without the library, Sonnet 5's code crashed: "Chart is not defined." Apparently, Sonnet 5 had written the page so that everything, even its plain text tables, ran after that step, and it couldn't catch the mistake.

The cost per run was essentially identical: $0.56 for Sonnet 5.5 and $0.57 for Sonnet 5. Also, Sonnet 5.5 finished the work faster.
So is Sonnet 5.5 a real step up? On this task, yes. Sonnet 5.5 shipped a dashboard that works and warns you about thin data, while Sonnet 5 shipped a blank page.
Claude Sonnet 5.5 Pricing and Availability
Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens on the API, unchanged from Sonnet 5 and half the price of Opus 5.5. Anthropic reports fewer tokens per task, which would lower the effective cost of a job even though the rate is flat. In our hands-on test, though, both models cost about the same to finish the task.
| Token type | Price per 1M tokens |
|---|---|
| Input | $2 |
| Output | $10 |
| Cache write (5-minute) | $2.50 |
| Cache write (1-hour) | $4 |
| Cache read | $0.20 |
The Batch API takes 50% off both input and output. The minimum cacheable prompt is 512 tokens, and on the Batch API the model can return up to 300K output tokens with the output-300k-2026-03-24 beta header.
Sonnet 5.5 is generally available, not a preview, with zero data retention on offer. Anthropic commits to not retiring it before September 28, 2027.
How to Get Access to Claude Sonnet 5.5?
You can use Claude Sonnet 5.5 today in the Claude apps, in Claude Code, and through the API with the model ID claude-sonnet-5-5. It's also available on Amazon Bedrock as anthropic.claude-sonnet-5-5, and on Google Cloud, Microsoft Foundry, and Claude Platform on AWS under claude-sonnet-5-5.
Here's a minimal Python call. Adaptive thinking is on by default, so the response can include thinking blocks ahead of the text:
from anthropic import Anthropic
client = Anthropic()
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=2048,
messages=[{"role": "user", "content": "Find the off-by-one bug in this loop: ..."}],
)
for block in response.content:
if block.type == "text":
print(block.text)
For a reak walkthrough of keys, costs, and your first project, see our complete guide to the Claude API. If you're building agents, our tutorial on building Claude Managed Agents is a good next step. We will publish a Claude Sonnet 5.5 API tutorial shortly.
Final Thoughts
Anthropic is betting that most paying work, including most coding-agent work, runs on this new mid-tier model. That also puts direct pressure on OpenAI's GPT-6 Sol, which trails it on Anthropic's knowledge-work numbers.
I'd switch now if you're on Sonnet 5, but do update yourself on the API changes. Keep Opus 5.5 for projects where the model has to decide what to build.
If you're keen to learn more about building with Anthropic's models, I recommend checking out our Introduction to Claude Models course.
FAQs
How does Claude Sonnet 5.5 compare to Claude Opus 5.5?
Claude Sonnet 5.5 scores higher than Opus 5.5 on Terminal-Bench 4.0 (70.6% vs 66.4%) and lands within a few points of it on GDPval-AA, OSWorld 2.1, and Humanity's Last Exam. Opus 5.5 costs twice as much per token and remains stronger on complex, open-ended work that needs sustained judgment.
How much does Claude Sonnet 5.5 cost?
Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5. Cache reads cost $0.20 per million tokens, and the Batch API takes 50% off. Anthropic says it costs up to 30% less per task than Sonnet 5 because it uses fewer tokens.
Where can I access Claude Sonnet 5.5?
Claude Sonnet 5.5 is available in the Claude apps, Claude Code, and the Claude API under the model ID claude-sonnet-5-5. It is also on Amazon Bedrock (anthropic.claude-sonnet-5-5), Google Cloud, Microsoft Foundry, and Claude Platform on AWS.
What is the context window of Claude Sonnet 5.5?
Claude Sonnet 5.5 has a 1M-token context window and a 128K-token maximum output on the Messages API. On the Message Batches API, it can return up to 300K output tokens with the output-300k-2026-03-24 beta header.
What changed in the Claude Sonnet 5.5 API compared to Sonnet 5?
Anthropic lists five breaking changes: adaptive thinking is on by default (use between_tools to turn off up-front thinking), forced tool use returns an error, thinking blocks are tied to the model and conversation, the computer_20251124 tool is rejected on the Claude API and Google Cloud, and the advisor tool rejects older models as advisors. Non-default temperature, top_p, or top_k values also return a 400 error.
