courses
Anthropic released Claude Sonnet 5.5 on September 28, 2026, and OpenAI answered a day later at DevDay with GPT-6.1 Sol.
Both are mid-tier models at the same list price, and both launch posts measure themselves against Claude Opus 5.5 rather than against each other. That leaves anyone picking between them with two sets of benchmarks that barely overlap and a price tag that doesn't break the tie.
In this article, I'll compare GPT-6.1 Sol and Claude Sonnet 5.5 across coding, knowledge work, speed, safety, and pricing, and then run both through the same hands-on coding test.
For deeper coverage of each model individually, see our GPT-6.1 Sol guide and our Claude Sonnet 5.5 guide.
TL;DR
- Both models charge the same per-token rates, so the bill only splits on cached context and very long prompts.
- GPT-6.1 Sol is cheaper for cache-heavy agent loops; Claude Sonnet 5.5 is far cheaper for prompts over 272K tokens.
- The vendors published almost no shared benchmarks, and each claims near-Opus 5.5 results on its own tests.
- In our hands-on test, Sol built the more careful app, but Sonnet got there faster and in fewer turns.
- Choose Sonnet 5.5 for speed, long context, and deployment on AWS or Google Cloud.
- Choose GPT-6.1 Sol for cache-heavy agents and pipelines that must reject bad inputs rather than caveat them.
What Is GPT-6.1 Sol?
GPT-6.1 Sol is OpenAI's mid-tier GPT-6 model, released at DevDay on September 29, 2026, just a week after GPT-6 Sol. OpenAI pitches it as "near-Astra intelligence for a fifth of the price" for agentic coding, computer use, and professional work.
To see the Sol line on a real project, try our GPT-6 Sol API tutorial, where you'll build a codebase migration agent, or our GPT-6 Sol vs Claude Opus 5.5 comparison.
What Is Claude Sonnet 5.5?
Claude Sonnet 5.5 is Anthropic's mid-tier model, released on September 28, 2026, as the second Claude 5.5 model and a faster, lower-cost complement to Claude Opus 5.5.
Anthropic says it is strongest at well-scoped everyday tasks, bug fixes, and polished documents, and early testers singled out its eye for design.
Our complete guide to the Claude API covers setup and cost management, and our Claude Opus 5.5 guide explains where its bigger sibling still earns its price.
GPT-6.1 Sol vs Claude Sonnet 5.5: Head-to-Head Comparison

GPT-6.1 Sol and Claude Sonnet 5.5 are closer than any recent cross-vendor pairing: same base rates, the same effort ladder, and near-identical context windows.
The differences that matter are in cache pricing, long-prompt billing, speed, and how each model handles inputs that break the rules.
| Feature | GPT-6.1 Sol | Claude Sonnet 5.5 |
|---|---|---|
| Developer and release | OpenAI, September 29, 2026 | Anthropic, September 28, 2026 |
| Positioning | Near-Astra results at a fifth of Astra's price | Faster, lower-cost complement to Opus 5.5 |
| Input / output per 1M tokens | $2 / $10 | $2 / $10 |
| Cached input read per 1M tokens | $0.10 | $0.20 |
| Long prompts | 2x input, 1.5x output above 272K input tokens | Standard rates across the full 1M window |
| Context window / max output | 1,050,000 / 128,000 tokens | 1M / 128K tokens |
| Reasoning effort levels | low, medium, high, xhigh, max | low, medium, high, xhigh, max |
| Artificial Analysis Intelligence Index | 51.8 | 56 |
| Major clouds | Microsoft Azure (Bedrock and Vertex AI not confirmed) | Amazon Bedrock, Google Cloud, Microsoft Foundry |
| Our hands-on test (rubric mean) | 5.0 | 4.7 |
Coding and agentic workflows
Neither vendor published a coding benchmark that the other also ran, so the only way to line them up is through Claude Opus 5.5, which both launch posts use as a yardstick.
Read that way, the two models land in roughly the same place: each claims Opus-level agentic results at half of Opus 5.5's $4 / $20 list price.
| Benchmark | GPT-6.1 Sol | Claude Sonnet 5.5 | Notes |
|---|---|---|---|
| DeepSWE v1.1 | Matches GPT-6 Astra; 6.4 points above GPT-6 Sol's best | Not published | OpenAI, relative claim only |
| AutomationBench 1.0.6 | 2.2 points above Opus 5.5 at medium effort | Not published | OpenAI, at roughly a third of Opus 5.5's cost |
| Terminal-Bench Science 0.1 | $5.47 per task at max effort vs $23.21 for Opus 5.5 | Not published | OpenAI; GPT-6 Astra leads the score at 68.1% |
| Terminal-Bench 4.0 | Not published | 70.6% (Opus 5.5: 66.4%) | Anthropic; Artificial Analysis measured 64% |
| CursorBench 4.0 | Not published | 55.5% (Opus 5.5: 57.8%) | Anthropic; tasks from real Cursor sessions |
| FrontierCode 1.1 (Main) | Not published | 52.1% at xhigh, 46.2% at max | Anthropic; Opus 5.5 scores 54.4% |
Sol's evidence is about long-horizon software engineering at low cost per task, while Sonnet's is about terminal work and messy, multi-file editor sessions, where it beats Opus 5.5 on one test and trails it on the other.
Two caveats apply to the whole table:
- Every figure is vendor-reported, with competitor results taken from public reports.
- OpenAI states most of its coding results relative to other models rather than as absolute scores.
On published numbers, coding is closer to a tie than either launch post suggests, which is why the hands-on test below carries more weight than usual.
Knowledge work and computer use
Claude Sonnet 5.5 has the stronger published case for office and document work, partly because Anthropic reported more of it.
Sonnet scores 1844 on GDPval-AA, within 2 points of Opus 5.5, and 80.1% on OSWorld 2.1. GPT-6.1 Sol beats Opus 5.5 with fallbacks on GDP.pdf, a question-answering test over complex professional PDFs, at less than half the cost per task, and comes within 2.1 points of GPT-6 Astra on the OSWorld 2.0 offline set.
The OSWorld versions differ, so those two results don't compare directly.
The one shared third-party figure, the Artificial Analysis Intelligence Index, puts Sonnet at 56 and Sol at 51.8, as reported by Smartscope and OpenRouter's model listing, respectively.
On the available evidence, Sonnet edges knowledge work.
Speed and token efficiency
Claude Sonnet 5.5 is the faster model by default.
Anthropic says it generates output 30%+ faster than Sonnet 5 and costs up to 30% less per task, and customer reports in the launch post back the pattern, such as Lovable's "a third fewer tool calls" on its coding evals.
GPT-6.1 Sol treats speed as a paid add-on. Its fast mode bills at 2x standard rates and isn't available with EU data residency, and OpenAI says GPT-6.1 Sol Ultrafast, with up to 8x faster token generation in Codex, arrives in the coming days.
Anthropic offers fast mode only on its Opus models, not on Sonnet 5.5.
If you want speed without paying a premium for it, Sonnet is the default.
Safety guardrails and reliability
Both models ship with guardrails that change behavior in production, and Sonnet's are the more visible ones.
Higher-risk cybersecurity requests on Sonnet 5.5 visibly fall back to Sonnet 5, and its thinking blocks only work in the account that produced them, which matters if you move conversations between accounts.
OpenAI rates GPT-6.1 Sol as Critical in cybersecurity.
It cuts factual errors at low effort from 11.4% to 7.7% against GPT-6 Sol, but its coding deception rate of 1.50% sits above Astra's 0.51%.
I'd keep a human review step on unattended Sol runs, and plan for the Sonnet 5 fallback if you do security work on Sonnet.
Pricing: what you actually pay
GPT-6.1 Sol and Claude Sonnet 5.5 charge the same $2 per 1M input tokens and $10 per 1M output tokens, so the bill only diverges on cached context and very long prompts.

Sol wins when the same context gets reused over and over, and Sonnet wins as soon as a single request goes past 272K input tokens.
Token rates side by side
| Rate | GPT-6.1 Sol | Claude Sonnet 5.5 |
|---|---|---|
| Input, per 1M tokens | $2.00 | $2.00 |
| Output, per 1M tokens | $10.00 | $10.00 |
| Cached input read, per 1M tokens | $0.10 | $0.20 |
| Cache write, per 1M tokens | $2.50 | $2.50 (5-minute), $4 (1-hour) |
| Requests over 272K input tokens | 2x input and cache rates, 1.5x output, for the whole request | Standard rates across the full 1M window |
| Batch discount | 50% (Batch and Flex) | 50% (Batch API) |
| Fast mode | 2x standard rates | Not offered |
The surcharge is the line to watch: on Sol, one oversized retrieval step reprices the entire request, not just the overflow.
What a real workload costs
| Workload (monthly) | GPT-6.1 Sol | Claude Sonnet 5.5 | Difference |
|---|---|---|---|
| Balanced assistant: 1M in / 250K out | $4.50 | $4.50 | $0 (0%) |
| Retrieval, each request under 272K: 10M in / 1M out | $30 | $30 | $0 (0%) |
| Retrieval, each request over 272K: 10M in / 1M out | $55 | $30 | Sol costs $25 more (+83%) |
| Cache-heavy loop: 100K prefix written once, 1,000 cached reads, plus 5K fresh in / 1K out per request | $30.25 | $40.25 | Sonnet costs $10 more (+33%) |
Each total is (volume ÷ 1M) × rate, summed across input, output, and cache lines, at standard rates.
- Balanced and sub-threshold retrieval: identical rates give identical bills.
- Over-threshold retrieval: the same 10M tokens, split into requests above 272K, costs Sol $4 per 1M input and $15 per 1M output, and that alone turns a tie into an 83% premium.
- Cache-heavy loop: 100M cached-read tokens cost $10 on Sol and $20 on Sonnet, assuming reads land inside the cache window. The 105K-token prompt stays under Sol's threshold.
Neither vendor has published token counts for a shared workload, and we haven't measured how each tokenizer splits the same text, so treat these totals as a rate comparison rather than a bill.
How GPT-6.1 Sol and Claude Sonnet 5.5 Performed
Because the published benchmarks don't overlap, I ran one build task on both models myself, with the same prompt and the same tooling.
The test
The task was to build a single-file web page that animates Dijkstra's shortest-path algorithm on three small weighted graphs supplied in a graph.json file.
The hard part isn't the animation. It's the data: the prompt never mentions that the second graph cuts the target off from the source, or that the third contains a negative edge weight, which breaks the rules Dijkstra's algorithm depends on.
With no shared coding benchmark between these two models, this is the one direct comparison in the article.
Both models ran inside the same coding agent, OpenCode, at high reasoning effort, with one attempt each and no retries.
Each got an empty workspace holding only the graph file and the same five tool permissions: read, list, glob, grep, and edit, where edit covers writing files.
The agent config denied shell and network access, so neither model could open its own page to check it. Here's the prompt, verbatim:
You have a file, `graph.json`, in the working directory. It holds three scenarios. Each scenario is a small weighted undirected graph — 12 nodes with `x`/`y` positions on a 1200×800 plane, 20 to 24 edges with integer weights, and a `source` and a `target` node.
Build a single-file HTML page (inline CSS and JS, canvas or SVG, no build step, no external dependencies, no network) that visualizes Dijkstra's algorithm on those graphs. It needs:
- read `graph.json` and embed its contents in the page (a JS object or `<script>` block) so it works when opened as a file — do not `fetch` it;
- a scenario selector: the first scenario is selected on load, and choosing another one restarts the visualization on that graph;
- draw the selected graph at the given positions with every edge weight labelled;
- auto-start on load from `source` to `target` with no click required, animating one edge examination every 300 milliseconds: pop the closest unsettled node, then examine each of its edges one at a time — highlight the edge being examined, and update the neighbour's tentative distance label if it improves;
- three visibly distinct node states at all times: settled, frontier (reached but not settled), and unvisited; each reached node shows its current tentative distance;
- stop when the target is settled, highlight the shortest path from source to target, and show a table of the final distance from the source to every node;
- pause / resume, single-step, and restart controls;
- everything fits in one viewport of about 1440×900 without scrolling.
Ship it as one working file named `index.html`. Do not install packages. Do not open, screenshot, or headless-render the page (no Playwright, Puppeteer, or Chrome). Do not ask me clarifying questions — make reasonable assumptions and note them briefly in a comment at the top of the file.
I graded each page on three questions:
- Algorithm correctness: Is it really Dijkstra, with correct final distances, the right path, and nodes settling in order, and does it notice when the graph breaks the algorithm's rules?
- Visualization fidelity: Can you follow the algorithm from the picture alone?
- Controls and layout: Do the controls work, and does it fit on one screen?
What GPT-6.1 Sol produced
GPT-6.1 Sol built a polished page with a live notebook beside the graph: the current edge check written out as arithmetic, a frontier list sorted closest-first, and a distance table that updates as it goes. On the first graph, it found the correct path and distance, and stepping through by hand settled the nodes in exactly the order the answer key lists.

It also handled both traps. On the cut-off graph, it reported "No path to the target" and left the two unreachable nodes at infinity. On the negative-edge graph, it refused to run, highlighted the offending edge in red, and explained that there's no finite shortest path.
It scored 5 on all three questions.
What Claude Sonnet 5.5 produced
Claude Sonnet 5.5 built a plainer page with a running log that explains every edge check, and it was correct on the first graph, with the same path, distances, and settle order.
It also spotted both traps without being told, and even flagged that the second graph has 17 edges when the prompt promised 20 to 24.

The miss came on the negative-edge graph.
Sonnet showed a clear red warning naming the negative edge and saying Dijkstra's guarantee no longer holds, but then drew a green "shortest path" anyway and reported a distance as the answer.
That cost it a point on algorithm correctness, and it scored 5 on the other two questions.
Results
Turns count the model's round trips to finish the build, and tool calls count every file read, search, and write.
| Measure | GPT-6.1 Sol | Claude Sonnet 5.5 |
|---|---|---|
| Turns | 5 | 4 |
| Tool calls | 7 | 3 |
| Runnability | pass | pass |
| Algorithm correctness | 5 | 4 |
| Visualization fidelity | 5 | 5 |
| Controls and layout | 5 | 5 |
| Rubric score (mean of three axes) | 5.0 | 4.7 |
GPT-6.1 Sol was the more rigorous builder, and the negative-edge graph is the whole difference: it refused a graph that breaks the algorithm, where Sonnet warned and then answered anyway.
I'd still pick Claude Sonnet 5.5 for everyday build work, because it reached a correct, working page in fewer turns and tool calls and finished much faster.
The caveat is real, though. If your pipeline has to reject invalid input rather than caveat it, Sol's behavior is the one you want.
This is one run per model on one task, so it probes careful algorithm implementation, not coding ability in general, and it says nothing about token use or cost.
When to Choose GPT-6.1 Sol vs Claude Sonnet 5.5

With identical base rates, the choice comes down to how your prompts are shaped, where you deploy, and how strictly the model must refuse bad inputs.
Choose GPT-6.1 Sol if...
- Your agents reuse a large cached context. Cache reads cost half of Sonnet's rate, which makes the cache-heavy example in our pricing table $10 a month cheaper on Sol.
- Your pipeline must refuse invalid input. In our test, Sol stopped on a graph that broke the algorithm's rules instead of producing an answer under a warning.
- You automate multi-step business workflows. OpenAI reports Sol 2.2 points above Opus 5.5 on AutomationBench at medium effort, at roughly a third of the cost.
- You already work in ChatGPT Work, Codex, or Azure. Sol is available in all three, so there's no new platform to adopt.
Choose Claude Sonnet 5.5 if...
- Your prompts run past 272K tokens. Sonnet bills its full 1M window at standard rates, while Sol's surcharge makes the over-threshold retrieval example in our pricing table 83% more expensive.
- You iterate quickly on builds. Sonnet finished our test in fewer turns and tool calls, and Anthropic reports its fastest Sonnet output to date.
- You deploy on AWS or Google Cloud. Sonnet is on Amazon Bedrock, Google Cloud, and Microsoft Foundry, while we could only confirm Sol on Azure.
- Your work is terminal-heavy coding or polished documents. Those are the areas Anthropic's own benchmarks and early testers emphasize.
How to Get Started With GPT-6.1 Sol and Claude Sonnet 5.5
Both models are generally available through their vendors' APIs and through OpenRouter, but their cloud reach differs.
| Surface | GPT-6.1 Sol | Claude Sonnet 5.5 |
|---|---|---|
| Consumer app | ChatGPT Work (Plus, Pro, Business, Enterprise, Edu); not yet in Chat | Claude apps |
| First-party API | OpenAI API (Responses, Chat Completions without tools, Batch) | Claude API |
| Cloud platforms | Microsoft Azure; Bedrock and Vertex AI not found in sources | Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry |
| Coding agents | Codex, GitHub Copilot | Claude Code, GitHub Copilot |
| Third-party routers | OpenRouter | OpenRouter |
| API model ID | gpt-6.1-sol |
claude-sonnet-5-5 (Bedrock: anthropic.claude-sonnet-5-5) |
The biggest difference is cloud reach: Sonnet 5.5 runs on all three major clouds, while we could only confirm GPT-6.1 Sol on Azure. On the vendors' own APIs, the IDs are gpt-6.1-sol and claude-sonnet-5-5.
Making your first API call
The quickest way to try both is through OpenRouter, where they're a one-string swap. On OpenAI's own API, remember that Sol's tool calling needs the Responses API.
from openai import OpenAI
client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key="YOUR_OPENROUTER_KEY")
response = client.chat.completions.create(
model="openai/gpt-6.1-sol", # swap for "anthropic/claude-sonnet-5.5"
messages=[{"role": "user", "content": "Explain why Dijkstra's algorithm fails on negative edge weights."}],
)
print(response.choices[0].message.content)
For full setup on each vendor's own API, see our GPT-6 Sol API tutorial and our complete guide to the Claude API.
Final Thoughts
GPT-6.1 Sol and Claude Sonnet 5.5 cost the same per token, so pick on workload shape.
Use Sonnet 5.5 for long prompts, fast iteration, and multi-cloud deployment, and GPT-6.1 Sol for cache-heavy agents and pipelines that must reject bad inputs.
My default for everyday coding would be Sonnet, with the caveat from our test: it saw that a graph broke Dijkstra's rules, said so, and answered anyway, where Sol refused.
What I find most telling is that both vendors benchmarked against Claude Opus 5.5 rather than each other, a sign they both expect most paid work to run on the mid tier.
If you're keen to build with models like these, I recommend our Associate AI Engineer for Developers track, or the Introduction to Claude Models and Working with the OpenAI API courses once you've picked one.
FAQs
Is GPT-6.1 Sol or Claude Sonnet 5.5 better for coding?
On published benchmarks it is close to a tie, because OpenAI and Anthropic reported different coding tests and each measured against Claude Opus 5.5. In our hands-on Dijkstra visualizer test, GPT-6.1 Sol was more rigorous and refused a graph with a negative edge weight, while Claude Sonnet 5.5 warned but still answered. Sonnet finished in fewer turns and tool calls, so it is the better pick for fast iteration and Sol for pipelines that must reject bad inputs.
How much do GPT-6.1 Sol and Claude Sonnet 5.5 cost?
Both cost $2 per 1M input tokens and $10 per 1M output tokens. GPT-6.1 Sol charges $0.10 per 1M cached input tokens against Sonnet 5.5's $0.20, but Sol bills any request over 272K input tokens at 2x input and 1.5x output. Claude Sonnet 5.5 bills its full 1M-token window at standard rates.
What are the API model IDs for GPT-6.1 Sol and Claude Sonnet 5.5?
GPT-6.1 Sol is gpt-6.1-sol on the OpenAI API, and Claude Sonnet 5.5 is claude-sonnet-5-5 on the Claude API, or anthropic.claude-sonnet-5-5 on Amazon Bedrock. On OpenRouter they are openai/gpt-6.1-sol and anthropic/claude-sonnet-5.5.
Where can I access GPT-6.1 Sol and Claude Sonnet 5.5?
GPT-6.1 Sol is in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, the OpenAI API, Microsoft Azure, GitHub Copilot, and OpenRouter. Claude Sonnet 5.5 is in the Claude apps, Claude Code, the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, GitHub Copilot, and OpenRouter.
Which has the larger context window, GPT-6.1 Sol or Claude Sonnet 5.5?
They are nearly identical: GPT-6.1 Sol takes 1,050,000 tokens and Claude Sonnet 5.5 takes 1M, and both return up to 128,000 output tokens. The difference is billing, because GPT-6.1 Sol reprices requests over 272K input tokens while Sonnet 5.5 does not.
A senior editor in the AI and edtech space. Committed to exploring data and AI trends.
