OpenAI used DevDay on September 29 to ship GPT-6.1 Sol, an upgrade to its mid-tier model that it says comes close to flagship GPT-6 Astra at a fifth of Astra's token prices. A week earlier, Anthropic launched Claude Opus 5.5, pitched as Fable 5.1-level performance for 40% less.
These are efficiency plays, and OpenAI's launch post names Opus 5.5 directly, claiming wins on business workflows and document work at a fraction of the cost per task. OpenAI's own charts tell a more interesting story than its headline, though, because the comparison changes depending on which effort setting you read.
Now, in this article, we're comparing GPT-6.1 Sol with Opus 5.5 largely because OpenAI's launch post makes that matchup directly, but Claude Sonnet 5.5 is another fair comparison, which we cover in our GPT-6.1 Sol vs Claude Sonnet 5.5 article.
TL;DR
- GPT-6.1 Sol is the cost-per-task leader: on every benchmark both models share, it gets there for a fraction of Opus 5.5's spend.
- However, the price gap mostly disappears on prompts over 272K tokens, where GPT-6.1 Sol's surcharge kicks in.
- Opus 5.5 still has the higher ceiling on business automation and scientific terminal work when you run it at max effort.
What Is GPT-6.1 Sol?
GPT-6.1 Sol is OpenAI's mid-tier reasoning model in the GPT-6 family, released on September 29, 2026 as an upgrade to GPT-6 Sol.
Its selling point is near-Astra results on coding, computer use, and professional work at a much lower price, with cheaper cached input aimed at agents that reuse context across requests.
Our GPT-6.1 Sol guide covers the launch; our GPT-6 Sol API tutorial builds a codebase migration agent on the previous version.
What Is Claude Opus 5.5?
Claude Opus 5.5 is Anthropic's flagship Opus model, released as the first model in the Claude 5.5 family.
Anthropic positions it for long-running agentic coding and knowledge work, matching Claude Fable 5.1 on most tasks at a lower cost, with adaptive thinking that is always on.
Our Claude Opus 5.5 guide covers the launch; our Opus 5.5 API tutorial builds an AI incident investigator.
GPT-6.1 Sol vs Claude Opus 5.5: Head-to-Head Comparison
GPT-6.1 Sol and Opus 5.5 overlap on only three published benchmarks, all from OpenAI's launch charts. Opus 5.5 posts the higher best score on two of them, and GPT-6.1 Sol reaches its results for 2x to 5x less per task on all three.
| Feature | GPT-6.1 Sol | Claude Opus 5.5 |
|---|---|---|
| AutomationBench (best score, cost per task) | 36.1% at max, $0.30 | 42.5% at max, $1.44 |
| GDP.pdf (best score, cost per task) | 32.0% at high, $0.35 | 28.8% at high, $0.83 |
| Terminal-Bench Science 0.1 (max effort) | 57.0%, $5.47 | 63.3%, $23.21 |
| DeepSWE v1.1 | 75.2% at high | Not published |
| Terminal-Bench 4.0 | Not published | 66.4% at max |
| Computer use | 71.4% on OSWorld 2.0 offline (max) | 81.8% on OSWorld 2.1 |
| Context window / max output | 1.05M / 128K | 1M / 128K |
| Knowledge cutoff | April 2026 | June 2026 |
| Effort levels | low, medium (default), high, xhigh, max | low, medium (default), high, xhigh, max |
The first three rows come from OpenAI's charts, which label the Claude results "Opus 5.5 w/ fallbacks". Everything else is each vendor reporting on its own model.
Business workflows and documents
This is where OpenAI picked its fight, and where the effort setting decides the winner.
OpenAI's headline says GPT-6.1 Sol scores 2.2 points above Opus 5.5 on AutomationBench, Zapier's test of agents completing multi-step workflows across dozens of business tools. That's true at medium effort, where GPT-6.1 Sol scores 31.7% at $0.19 per task against Opus 5.5's 29.5% at $0.65.
Turn both up to max, and the order flips. Opus 5.5 climbs to 42.5%, above even GPT-6 Astra, while GPT-6.1 Sol tops out at 36.1%. Opus 5.5 pays for that with nearly 5x the cost per task, so the question is whether six points of workflow success are worth it for your agent. For a support bot handling thousands of tickets, probably not. For a finance workflow where a failed run means a human redoes it, maybe.
GPT-6.1 Sol is the better value for document work and everyday automation. Opus 5.5 is the pick when the hardest workflow runs have to succeed.
Coding and scientific work
There's no shared coding benchmark, so each vendor's number stands alone in the table. OpenAI says GPT-6.1 Sol matches GPT-6 Astra on DeepSWE v1.1, which tests long-horizon engineering in real codebases, while Anthropic reports Opus 5.5 leading GPT-6 Astra on Terminal-Bench 4.0 and FrontierCode.
In our earlier hands-on test, GPT-6 Sol, the previous version, beat Opus 5.5 on a Tetris build at under a third of the run cost (see our GPT-6 Sol vs Claude Opus 5.5 comparison), so we suspect GPT-6.1 Sol wins out.
Computer use
The two models can't be compared here yet. OpenAI reports GPT-6.1 Sol on the offline set of OSWorld 2.0, within 2.1 points of GPT-6 Astra, while Anthropic reports Opus 5.5 on OSWorld 2.1, a newer version with different tasks. Neither vendor published the other's version, so the summary table's numbers aren't a head-to-head. Treat computer use as open until someone runs both on the same set.
Guardrails and API constraints
Opus 5.5 is the more restricted model in practice. Anthropic routes most cybersecurity tasks to Opus 4.8 unless you're in its Cyber Verification Program, and biology work goes through similar fallbacks. That's also why OpenAI's charts label its Claude results "with fallbacks".
GPT-6.1 Sol's constraints are about the API instead. It has no none or minimal effort setting, and tool calling works only through the Responses API, not Chat Completions. Opus 5.5 keeps adaptive thinking always on and rejects forced tool use. For security teams, Opus 5.5's routing is the bigger practical difference; for developers migrating code, it's GPT-6.1 Sol's API changes.
Pricing: what you actually pay
GPT-6.1 Sol costs exactly half of Opus 5.5 on every standard rate, until a prompt passes 272K input tokens.
Token rates side by side
| Rate | GPT-6.1 Sol | Claude Opus 5.5 |
|---|---|---|
| Input, per 1M tokens | $2.00 | $4.00 |
| Output, per 1M tokens | $10.00 | $20.00 |
| Cached input read, per 1M tokens | $0.10 | $0.20 |
| Cache write, per 1M tokens | $2.50 | $5.00 (5-minute), $8.00 (1-hour) |
| Long-context surcharge | Over 272K input tokens: 2x input and cache, 1.5x output for the whole request | None |
| Batch discount | 50% | 50% |
| Fast mode | 2x standard rates | $8 / $40 per 1M tokens |
The flat 0.5x multiple holds across input, output, and cache reads, so for normal-sized prompts the comparison collapses to "GPT-6.1 Sol costs half". Both models bill reasoning tokens as output, which matters more for Opus 5.5 since its thinking can't be switched off.
What a real workload costs
| Workload | GPT-6.1 Sol | Claude Opus 5.5 | Difference |
|---|---|---|---|
| Balanced assistant: 1M in / 250K out | $4.50 | $9.00 | $4.50 (50% less) |
| Retrieval, each request under 272K: 10M in / 1M out | $30 | $60 | $30 (50% less) |
| Retrieval, each request over 272K: 10M in / 1M out | $55 | $60 | $5 (8% less) |
Each total is (volume ÷ 1M) × rate, summed across input and output, at standard rates.
- Balanced assistant: the plain half-price case, and the one most chat and agent workloads look like.
- Retrieval under the threshold: still half price, so RAG setups that keep each prompt small get the full saving.
- Retrieval over the threshold: the surcharge doubles GPT-6.1 Sol's input rate to Opus 5.5's level, and the saving shrinks to a rounding error. If your prompts routinely carry whole codebases, budget as if the two cost the same.
The two vendors use different tokenizers, and neither has published token counts for a shared workload, so treat these totals as a rate comparison rather than a bill.
How GPT-6.1 Sol and Claude Opus 5.5 Performed
The benchmarks above are the vendors' own, so I ran one build task against GPT-6.1 Sol and Claude Opus 5.5 with an identical prompt and identical tooling.
The task: a single-file HTML appointment scheduler for a community health clinic, with a Monday-to-Sunday week view of physio, dental, imaging, and pediatrics appointments, a form to add, edit, and delete bookings, localStorage persistence, and a seeded sample week of about ten appointments. No build step, no dependencies, no network.
The hard part is the conflict logic, which has to hold several rules at once:
- Two appointments clash if they overlap in the same room, or if they overlap with the same clinician
- Back-to-back bookings, where one ends exactly when the next starts, must not clash
- Every flag has to name the other appointment and the reason, not just turn red
- Flags have to update after every edit and delete, and survive a reload
- The seeded week has to contain exactly two room clashes and one clinician clash
I like this test because it probes OpenAI's claim that GPT-6.1 Sol matches GPT-6 Astra on real-codebase engineering and Anthropic's pitch for Opus 5.5 as a long-running coding agent.
Here's GPT-6.1 Sol's scheduler after I moved a booking, deleted one, and added a new one that double-books a clinician. The new Thursday booking is flagged, naming the appointment it clashes with:

And here's Opus 5.5's after the same steps, with its clash panel listing each remaining pair and how long the overlap lasts:

The main findings:
- On the conflict logic, nobody separated. Both pages opened with exactly two room clashes and one clinician clash flagged, each naming the other appointment and the reason, and both left back-to-back bookings alone.
- Edits, deletes, and reloads caught nobody out either. Moving a booking to a free room cleared both halves of its clash, deleting one half of the clinician clash cleared the other, a new booking that double-booked a clinician was flagged correctly, and everything survived a reload.
- The differences were all presentation. GPT-6.1 Sol built the more polished page, with summary cards, a service filter, and a "Conflicts only" toggle, but it lists each day's bookings in time order rather than on a time grid. Opus 5.5 added a clash panel that shows how long each overlap lasts and warns you about clashes before you save, but its week grid needs a scroll.
- Opus 5.5 wrote the trickier sample week. Its seed included an overlap in different rooms with different clinicians.
I scored both 5 out of 5 for spec adherence and conflict logic. GPT-6.1 Sol got 5 for usability and layout, and Opus 5.5 got 4 because its week view doesn't fit on one screen.
The cost is where they split. Opus 5.5 spent more than twice as many output tokens, most of them on thinking:
- GPT-6.1 Sol: $0.27
- Claude Opus 5.5: $1.11
So which one wins? On this task, GPT-6.1 Sol, and mostly only on price. Both shipped a correct, working scheduler, and GPT-6.1 Sol did it for about a quarter of the cost.
When to Choose GPT-6.1 Sol vs Claude Opus 5.5
For most teams, the fork is cost per task: GPT-6.1 Sol reaches OpenAI's reported scores for roughly a fifth to a half of what Opus 5.5 spends. The exceptions are long prompts, the last few points of success on hard runs, and where you deploy.
Choose GPT-6.1 Sol if...
- You run agents at volume. On business-workflow automation, it beats Opus 5.5 at medium effort for about a third of the cost per task, which compounds across thousands of runs.
- Your work is questions over dense documents. It leads Opus 5.5 on OpenAI's PDF benchmark at every effort setting, at less than half the cost.
- Your team already works in ChatGPT Work or Codex. It's the model OpenAI recommends there for complex coding and agentic work.
- You reuse long system prompts or context. Cached input at $0.10 per million tokens makes repeat-heavy agent loops cheap, as long as each prompt stays under 272K tokens.
Choose Claude Opus 5.5 if...
- Your prompts regularly exceed 272K tokens. GPT-6.1 Sol's surcharge erases most of its price advantage there, and Opus 5.5 has no long-context surcharge.
- A failed run costs more than the tokens. At max effort, Opus 5.5 posts the highest AutomationBench score in OpenAI's own chart, including GPT-6 Astra.
- You deploy through Cursor, Amazon Bedrock, or Google Vertex AI. Opus 5.5 is listed on all three; GPT-6.1 Sol isn't listed in the Cursor or Vertex AI docs yet.
- You want Anthropic's conservative defaults for unattended agents. Its safeguards route risky security and biology work to other models rather than attempting it.
How to Get Started With GPT-6.1 Sol and Claude Opus 5.5
| Surface | GPT-6.1 Sol | Claude Opus 5.5 |
|---|---|---|
| Consumer app | ChatGPT Work and Codex (Plus, Pro, Business, Enterprise, Edu); not yet in Chat | Claude apps (Pro, Max, Team, Enterprise) |
| First-party API | OpenAI API (tool calling via the Responses API) | Claude API |
| Cloud platforms | Azure AI Foundry; not listed in the Vertex AI docs as of September 30, 2026 | Amazon Bedrock, Google Vertex AI, Azure AI Foundry |
| Coding agents | Codex, GitHub Copilot; not listed in the Cursor docs as of September 30, 2026 | Claude Code, Cursor, GitHub Copilot |
| Third-party routers | OpenRouter (openai/gpt-6.1-sol) | OpenRouter (anthropic/claude-opus-5.5) |
| API model ID | gpt-6.1-sol |
claude-opus-5-5 |
The biggest difference is reach: Opus 5.5 is on all three major clouds and in Cursor, while GPT-6.1 Sol is centered on OpenAI's own products, Azure, and Copilot.
Making your first API call
The two models use different SDKs, so switching means changing the client as well as the model string:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-6.1-sol",
reasoning={"effort": "medium"},
input="Summarize the payment terms in this contract: ...",
)
print(response.output_text)
from anthropic import Anthropic
client = Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=2048,
messages=[{"role": "user", "content": "Summarize the payment terms in this contract: ..."}],
)
print(next(b.text for b in response.content if b.type == "text"))
For full builds, see our GPT-6 Sol API tutorial and our Opus 5.5 API tutorial, which I also shared earlier.
Final Thoughts
If cost per task is your constraint, use GPT-6.1 Sol. If you need the highest success rate on hard agent runs, send long prompts, or deploy through Cursor or Bedrock, use Opus 5.5.
What I find most telling is that OpenAI's own charts make the case for both. It chose the medium-effort point for its headline, but the same chart shows Opus 5.5 on top at max effort. OpenAI is betting that most buyers care about the cheaper point on the curve, and for everyday agents, I think it's right.
Because Sol is a mid-tier model, GPT-6.1 Sol vs. Claude Sonnet 5.5 is worth testing, too.
And if you're building on either model, I recommend our OpenAI Fundamentals skill track or our Introduction to Claude Models course.
FAQs
Is GPT-6.1 Sol better than Claude Opus 5.5?
It depends on the effort setting you compare. In OpenAI's launch charts, GPT-6.1 Sol beats Opus 5.5 on AutomationBench at medium effort and on the GDP.pdf document benchmark at every setting, at a fraction of the cost per task. At max effort, Opus 5.5 scores higher on AutomationBench and Terminal-Bench Science 0.1, but costs roughly 4x to 5x more per task.
How much cheaper is GPT-6.1 Sol than Claude Opus 5.5?
GPT-6.1 Sol's API rates are exactly half of Opus 5.5's: $2 vs $4 per million input tokens and $10 vs $20 per million output tokens. The gap nearly closes on prompts over 272K input tokens, where GPT-6.1 Sol charges 2x input and 1.5x output for the whole request and Opus 5.5 has no surcharge.
Where can I use GPT-6.1 Sol and Claude Opus 5.5?
GPT-6.1 Sol is in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, the OpenAI API, Azure AI Foundry, GitHub Copilot, and OpenRouter. Opus 5.5 is in the Claude apps, Claude Code, the Claude API, Amazon Bedrock, Google Vertex AI, Azure AI Foundry, Cursor, GitHub Copilot, and OpenRouter.
What are the API model IDs for GPT-6.1 Sol and Claude Opus 5.5?
The OpenAI API model ID is gpt-6.1-sol, and the Claude API model ID is claude-opus-5-5. On OpenRouter they are openai/gpt-6.1-sol and anthropic/claude-opus-5.5.
Which is better for long documents, GPT-6.1 Sol or Claude Opus 5.5?
For questions over complex PDFs, GPT-6.1 Sol scores higher than Opus 5.5 on OpenAI's GDP.pdf benchmark at less than half the cost per task. For very long prompts over 272K tokens, Opus 5.5 is competitive on price because GPT-6.1 Sol's long-context surcharge brings the two close to parity.