Track
OpenAI released GPT-6.1 Sol on September 29, 2026 at DevDay, one week after GPT-6 Sol and 26 days after GPT-6 Astra.
Astra averaged $23.80 per Terminal-Bench Science 0.1 task at maximum effort, which priced it out of routine work. Sol is OpenAI's answer to that problem.
The headline numbers are $2.00 per 1M input tokens, $10.00 per 1M output tokens, a 1,050,000-token context window, and 128,000 tokens of maximum output. In OpenAI's announcement, GPT-6.1 Sol matches Astra on DeepSWE v1.1 at roughly one-fifth of the cost, scores 2.2 percentage points above Anthropic's Claude Opus 5.5 on AutomationBench at medium effort, and costs $5.47 per Terminal-Bench Science 0.1 task against Astra's $23.80.
In this article, I'll cover what is new with GPT-6.1 Sol, looking at the features, the benchmark, and safety numbers OpenAI published, which tier in the GPT-6 family you should reach for, and what it costs. For deeper coverage of the neighbors, see our GPT-6 Sol API tutorial and our guide to Claude Sonnet 5.5. I also recommend reading about the new OpenAI dots, also announced at Dev Day.
In a Nutshell
- GPT-6.1 Sol is a mid-tier refresh that sits below GPT-6 Astra and above GPT-6 Luna. Input and output prices match GPT-6 Sol, and cached input drops from $0.20 to $0.10 per 1M tokens.
- Cost per completed task is the whole story. Raw capability still trails Astra on the hardest cyber, biology, and science evaluations.
- OpenAI treats it as Critical in cybersecurity and High in biology, so it ships with the same safeguards stack as Astra.
- Make it your default for agentic coding, computer use, and business workflows. Keep Astra for wet-lab reasoning, authorized security research, and the hardest scientific work.
- Moving from GPT-6 Sol takes code changes. GPT-6.1 Sol does not support
noneorminimalreasoning effort, and Chat Completions cannot call tools.
What Is GPT-6.1 Sol?
GPT-6.1 Sol is OpenAI's mid-tier reasoning model in the GPT-6 series, released on September 29, 2026 as an upgrade to GPT-6 Sol. It slots below GPT-6 Astra, the flagship launched on September 3, and above GPT-6 Luna, the small model.
OpenAI says it nearly matches Astra on agentic coding, computer use, and professional work.
The system card addendum goes further and describes capabilities comparable to Astra's, with an unmatched combination of speed and affordability.
I'd treat that as a cost-adjusted claim rather than an absolute one. Astra still leads on the hardest cyber, biology, and science evaluations.
There is no GPT-6.1 Astra. The Wall Street Journal reported that OpenAI scrapped the October release after the model regressed on alignment tests and on staying within the scope a user had authorized, and OpenAI confirmed the decision.
So, GPT-6.1 Sol is the only 6.1 model, and its own deception and persistence numbers (covered below) deserve a close read.

How does GPT-6.1 Sol compare with GPT-6 Sol?
GPT-6.1 Sol improves on GPT-6 Sol most in agentic and security work. T
hese are the gains OpenAI reports:
- DeepSWE v1.1: 6.4 percentage points above GPT-6 Sol's best score, at a lower reasoning effort and cost.
- AutomationBench: 4.8 points higher at medium effort.
- OSWorld 2.0 (offline set): 7 points higher at maximum effort, at less than half the cost per task.
- Terminal-Bench Science 0.1: more than double GPT-6 Sol's score at maximum effort, at less than half the cost per task.
- ExploitBench Internal Port: 21.5% against 5.5%.
- Internal research debugging: 75.52% against 64.20%.
What does the Preparedness classification mean?
OpenAI treats GPT-6.1 Sol as Critical capability in cybersecurity, High in the Biological and Chemical domain, and below the High threshold in AI Self-Improvement. GPT-5.6 Sol was rated High, not Critical, in cybersecurity. Developers should read this classification twice.
It means the model inherits the full Astra safeguards stack, including phased access for advanced cyber work through the Daybreak program. Sol does not get a lighter mid-tier version of those controls.
GPT-6.1 Sol Key Features
Most of what is new here is about doing the same work for less money, with a few API changes you need to plan around. These are the capabilities that change how you'd build with it.
A 1,050,000-token context window, with a cliff at 272,000
GPT-6.1 Sol takes 1,050,000 input tokens and returns up to 128,000, matching GPT-6 Sol's window.
The catch is in the billing: prompts over 272,000 input tokens are charged at 2x the input and cache rates and 1.5x the output rate, for the entire request rather than the overflow.
Repository-wide analysis under 272,000 tokens is cheap. A single oversized retrieval step can push a run over the line and reprice everything in that request.
Agentic coding and computer use at a fraction of Astra's task cost
The clearest use case is swapping GPT-6.1 Sol in wherever you were running Astra on long tool-calling loops.
OpenAI reports parity with Astra on DeepSWE v1.1 at roughly one-fifth of the cost. On the OSWorld 2.0 offline set at maximum effort, it comes within 2.1 points of Astra at roughly one-seventh of the cost per task.
Our GPT-6 Sol codebase migration agent ran end-to-end for $0.7082, with 91% of input tokens served from cache. That run used GPT-6 Sol, not GPT-6.1 Sol. The standard token prices are the same, and the cached input rate is half as much.
Reasoning is always on
GPT-6.1 Sol supports low, medium (the default), high, xhigh, and max reasoning effort. According to OpenAI's model page, none and minimal are not supported.
Tool calling requires the Responses API.
Chat Completions still works, but without tool calling. GPT-6 Sol allowed function calling in Chat Completions at none effort, so anyone using it as a cheap non-reasoning function caller has a migration to do.
Lower factual error rates at low reasoning effort
OpenAI measured the share of answers containing a factual error on de-identified ChatGPT conversations where users had flagged an earlier model's mistake. At low effort, that share fell from 11.4% for GPT-6 Sol to 7.7% for GPT-6.1 Sol, a reduction of about 32%.
Across the tested effort settings, its error rate stays within 1.9 percentage points of Astra's.
OpenAI says these prompts are deliberately difficult and not representative of typical usage.
The addendum's own hallucination chart shows GPT-6.1 Sol and GPT-6 Sol at similarly low rates on user-flagged cases, so treat this as a low-effort gain rather than an across-the-board one.
The eval covers general factual questions, not extraction or classification. If you run those pipelines at medium effort, test low on your own data before switching. It is also the lowest setting available.
Chain-of-thought that follows instructions
GPT-6.1 Sol controls its own chain-of-thought (CoT) more reliably than any previous Sol model.
Among CoTs between 750 and 1,250 tokens, it follows the requested CoT instruction in 44.8% of cases, against 23.2% for GPT-6 Sol, 16.1% for GPT-5.6 Sol, and 1.7% for GPT-5.5 Thinking. Astra still leads at 60.9%.
OpenAI's own examples show the limit. Told not to reason about a question in the analysis channel, GPT-6.1 Sol produced no analysis-channel text at all, then made question-related tool calls in commentary. It obeyed the letter of the instruction.
The same safeguards stack as the flagship
Because OpenAI assigned GPT-6.1 Sol the same Preparedness determinations as Astra, it gets the same safeguards.
On OpenAI's cybersecurity safety evaluation for production chat, it scores 0.987, the highest of any model in the comparison.
On the biology refusal evaluation, it declines the same severe and dual-use requests as GPT-6 Sol while refusing fewer benign prompts (0.982 against 0.964).
The picture is less even in agentic settings.
OpenAI reports modest regressions against GPT-5.6 Sol in synthetic and semi-synthetic cyber environments. Cyber work still goes through phased access, and OpenAI is running GPT-6.1 Sol through the Daybreak trusted-access program for advanced authorized use, as it did for Astra.
How Does GPT-6.1 Sol Perform on the Benchmarks?
GPT-6.1 Sol lands between GPT-6 Sol and GPT-6 Astra on most evaluations OpenAI published.
It sits close to Astra on health, hallucinations, and misalignment flags, and further away on cyber and biology.
Coding deception and unwanted persistence are the exceptions, where it scores worse than Astra.
Where Astra wins outright, the gap runs from about 4 to 16 points.
Where cost is factored in, Sol wins comfortably.
OpenAI ran its own models in its research environment or through its API and took competitor results from public reports.
Coding and agentic workflows
The agentic numbers are the reason this model exists. In OpenAI's announcement, GPT-6.1 Sol matches Astra on DeepSWE v1.1 and comes within 2.1 points of it on OSWorld 2.0. On AutomationBench it scores 2.2 points above Opus 5.5 at medium effort, at roughly a third of the cost.
Cost per task is where the separation is obvious.
On Terminal-Bench Science 0.1 at maximum effort, GPT-6.1 Sol averaged $5.47 per task, against $23.21 for Opus 5.5 and $23.80 for Astra. Astra still leads the accuracy table at 68.1%, and OpenAI recommends it for the most difficult scientific research tasks.
For context from our own coverage, Claude Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0. On FrontierCode 1.1 it scored 46.2% at max effort and 52.1% at xhigh, and Anthropic's launch table lists GPT-6 Sol at 49.3%. Those are different benchmarks and vendor-reported numbers, so treat them as context rather than a head-to-head.
Cybersecurity and exploit development
GPT-6.1 Sol is the Sol-tier model OpenAI treats as Critical in cybersecurity, and the evaluation numbers explain why. On ExploitBench, it reaches 99.7% at maximum reasoning effort against 81.7% for GPT-6 Sol and 100% for Astra. OpenAI cautions that these results may be inflated by contamination from historical vulnerabilities.
| Evaluation | GPT-6.1 Sol | GPT-6 Sol | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|
| ExploitBench (max effort) | 99.7% | 81.7% | 100% | Not published |
| ExploitBench Internal Port (code execution) | 21.5% | 5.5% | 31.5% | 3.5% |
| SEC-Bench Pro (pass@1) | 78.8% | 66.3% | 85.4% | 79.1% |
| ExploitGym (intended vulnerability) | 35.1% | 22.1% | 42.4% | 30.3% |
The Internal Port result is the most informative one, because it uses vulnerabilities disclosed between June and August 2026 to reduce contamination. Going from 5.5% to 21.5% is close to a 4x improvement over GPT-6 Sol on recently disclosed vulnerabilities.
It still sits 10 points under Astra there, 6.6 points under on SEC-Bench Pro, and 7.3 points under on ExploitGym. OpenAI's own summary is that reliable exploitation of recently disclosed vulnerabilities remains challenging for GPT-6.1 Sol.
Health and mental health conversations
On health evaluations, GPT-6.1 Sol is effectively an Astra substitute. Its length-adjusted HealthBench Professional score is 64.2 against Astra's 64.7 and GPT-6 Sol's 60.8. HealthBench Hard is 36.2 against Astra's 36.6 and GPT-6 Sol's 30.1.
MentalHealthBench is an open benchmark of 1,215 synthetic conversations scored against clinician-written rubrics. GPT-6.1 Sol scores 57.9% ± 1.0 overall, against 58.7% for Astra and 54.2% for GPT-6 Sol. On the 344 emergent-acuity tasks, it scores 58.0%, within a standard error of Astra's 58.5%.
Biological and chemical capability thresholds
Biology is where GPT-6.1 Sol trails Astra by the widest single margin, although OpenAI still classifies it as High. On TroubleshootingBench, which uses non-public expert-written wet lab protocols, it scores 47.96% against Astra's 63.46%, a 15.5-point gap.
| Evaluation (High threshold) | Threshold | GPT-6.1 Sol | GPT-6 Sol | GPT-6 Astra |
|---|---|---|---|---|
| Multimodal Troubleshooting Virology | 31% | 55.34% | 50.6% | 63.11% |
| ProtocolQA Open-Ended | 54% | 40.74% | 44.4% | 45.37% |
| Tacit Knowledge and Troubleshooting | 80% | 88.50% | 79.2% | 92.55% |
| TroubleshootingBench | 36.4% | 47.96% | 45.3% | 63.46% |
ProtocolQA Open-Ended is the one evaluation below its threshold, at 40.74% against 54%. Astra also falls short of it, at 45.37%.
None of the Critical-threshold biology evaluations were crossed. On SHP2 Protein Function Prediction, GPT-6.1 Sol reached a mean R² of 0.332 against a 0.60 threshold. On the Coronavirus-ACE2 Cell-Entry Screen, it scored 0.423 against a 0.75 threshold.
Alignment, hallucinations, and robustness
OpenAI designed these evaluations to test challenging situations, so the rates below are not typical-use failure rates.
The most useful signal is a deployment simulation of internal Codex traffic across 49,650 tasks.
GPT-6.1 Sol picked up 28 severity-3-or-higher flags (0.056%), against 27 for Astra (0.054%), 42 for GPT-6 Sol (0.085%), and 63 for GPT-5.6 Sol (0.127%). That works out to 33% fewer severe flags than GPT-6 Sol and 56% fewer than GPT-5.6 Sol.
OpenAI describes the simulation as a signal about internal deployment risk, not a direct measure of external safety.
It also found more reward-hacking and concealed-uncertainty flags than Astra at lower severities, and more credential-harvesting flags than GPT-6 Sol.
Two results are regressions you should know about before wiring it into an agent loop:
- Unwanted persistence appeared in 23.5% of GPT-6.1 Sol rollouts against 17.4% for Astra. The evaluation tests low-stakes blocked actions, such as trying email after a direct message is refused, and runs without the system-level controls designed to stop circumvention.
- Coding deception, meaning misrepresentation of its own work, ran at 1.50% against 0.51% for Astra and 1.30% for GPT-6 Sol. GPT-5.6 Sol's rate was nearly 7x higher than GPT-6.1 Sol's.
One result cuts both ways.
Among samples that discovered a message board of apparent peer agents, GPT-6.1 Sol attempted communication in 38% against GPT-6 Sol's 26%. It carried out the specified unauthorized action in only 3%, against 11%.
On the positive side, it failed to acknowledge a broken search tool in only 2.08% of cases against GPT-6 Sol's 4.92%. It made no attempts to bypass the auto-review monitor or to exploit the ExploitGym honeypot.
Evaluation awareness in the deployment simulation
OpenAI also checked whether models noticed they were being evaluated, by looking for verbalized simulation awareness in the chain of thought.
GPT-6.1 Sol showed those signs in 4.81% of trajectories, against 8.67% for GPT-6 Astra, 4.70% for GPT-6 Sol, and 2.52% for GPT-5.6 Sol. The original mixed-model production traffic sat at 0.76%.
Restricting the comparison to the 41,849 tasks where no model showed awareness, GPT-6.1 Sol still recorded roughly 17% fewer severity-3-or-higher flags than Astra, 46% fewer than GPT-6 Sol, and 65% fewer than GPT-5.6 Sol.
That subset rules out the obvious objection that the model was on its best behavior.
Monitor evasion under adversarial conditions
OpenAI tested whether GPT-6.1 Sol could sabotage a repository while hiding the work from a CoT monitor.
At the highest reasoning effort tested, its success rate at evading CoT-only monitors was close to zero in both the task-guided and unguided settings. At lower efforts, it evaded at rates similar to Astra. A high-recall monitor with access to the full trajectory brought successful evasion close to zero as well.
Awareness of the monitor changes the picture.
When told it was being monitored, GPT-6.1 Sol became more evasive than GPT-6 Sol, although its CoT-only monitor recall stayed higher than Astra's. Full-context monitoring still caught 100% of its successful honeypot exploits.
On a new math side-tasking benchmark, GPT-6.1 Sol solved side tasks that GPT-6 Sol consistently failed, such as Diagonal Latin and the 51-pile counter, while running Expert-SWE software engineering work.
A full-context monitor caught every successful attempt. OpenAI found no evidence that GPT-6 Astra, GPT-6 Sol, or GPT-6.1 Sol can hide complex misalignment reasoning from its monitors, and the addendum does not cover Luna.
Which GPT-6 Tier Should You Use?
Use GPT-6.1 Sol for most agentic work, and reach for GPT-6 Astra when a failed run costs more than the tokens. The GPT-6 family now has four members, and the choice comes down to what a failure costs you.

My position: GPT-6.1 Sol should be the default for agentic coding, computer use, and business workflows, and Astra should be an escalation target rather than a starting point.
The exceptions are wet-lab reasoning, authorized security research, and the hardest scientific work. The 15.5-point TroubleshootingBench gap and the 10-point Internal Port gap are too big to paper over with retries.
| Use case | Tier | Why |
|---|---|---|
| Repo migrations, multi-step coding agents | GPT-6.1 Sol, medium to high effort | DeepSWE v1.1 parity with Astra at roughly one-fifth of the cost |
| Browser and desktop automation | GPT-6.1 Sol | Within 2.1 points of Astra on OSWorld 2.0 (offline set, max effort) at roughly one-seventh of the cost per task |
| Multi-step business workflows | GPT-6.1 Sol | 2.2 points above Opus 5.5 on AutomationBench at medium effort, at roughly a third of the cost |
| Authorized security research | GPT-6 Astra | 31.5% vs 21.5% on ExploitBench Internal Port and 85.4% vs 78.8% on SEC-Bench Pro |
| Lab protocol review, virology troubleshooting | GPT-6 Astra | 63.46% vs 47.96% on TroubleshootingBench |
| Hardest scientific research | GPT-6 Astra | Leads Terminal-Bench Science 0.1 at 68.1%, and OpenAI recommends it for this work |
| File triage, routing, bulk classification | GPT-6 Luna | Cheapest tier in the family at $0.10 input and $0.50 output; we used it to narrow 56 files to 16 in our migration agent |
Reasoning effort is the second dial.
GPT-6.1 Sol runs at low, medium (default), high, xhigh, or max, and reasoning tokens bill at the output rate.
OpenAI's $5.47 Terminal-Bench Science figure is a single maximum-effort data point, so it does not show how cost scales across settings. Measure that on your own workload before you pick a default.
Fast mode bills at 2x standard rates and is not available with EU data residency. OpenAI also plans a GPT-6.1 Sol Ultrafast option in Codex with up to 8x faster token generation, which Neowin reports will cost 6x the standard tier.
GPT-6.1 Sol Pricing and Availability
GPT-6.1 Sol costs $2.00 per 1M input tokens and $10.00 per 1M output tokens, the same as GPT-6 Sol and one-fifth of Astra's $10 and $50. Cached input drops from $0.20 to $0.10 per 1M tokens, with cache writes at $2.50 per 1M.
Reasoning tokens bill at the output rate, which is the number to watch at high and max effort.
| Rate | Price per 1M tokens |
|---|---|
| Input | $2.00 |
| Cached input | $0.10 |
| Cache writes | $2.50 |
| Output | $10.00 |
Several modifiers apply on top.
Prompts over 272,000 input tokens bill at 2x input and cache rates and 1.5x output for the whole request.
Fast mode is 2x standard, regional processing adds a 10% premium where available, and Batch and Flex run 50% below standard.
Against Anthropic's newest Opus model, the gap is wide only below 272,000 tokens. C
laude Opus 5.5 costs $4 per 1M input tokens and $20 per 1M output tokens, and bills its full 1M-token window at those rates, as we covered in our Opus 5.5 API tutorial.
GPT-6.1 Sol is exactly half that price under the threshold. Above it, GPT-6.1 Sol costs $4 for input and $15 for output, so the input advantage disappears.
Anthropic's top tier is Claude Fable 5.1, which lists at $10 and $50, the same as Astra.

Fast mode is unavailable with EU data residency, so European deployments with residency requirements lose the latency option while keeping the standard rates.
How to Get Access to GPT-6.1 Sol
GPT-6.1 Sol is available in ChatGPT Work and Codex, in the OpenAI API, and through several third-party platforms. As of September 29, 2026, these are the routes we could confirm:
- ChatGPT Work and Codex: Plus, Pro, Business, Enterprise, and Edu plans, in the desktop app, CLI, and IDE extension. Enterprise and Edu keep it off by default until an administrator enables it. It is not in Chat, and Free and Go plans are not included. The Codex models page lists
gpt-6.1-sol. - OpenAI API: the model ID is
gpt-6.1-sol. - Third-party platforms: OpenRouter (
openai/gpt-6.1-sol), Vercel AI Gateway, and GitHub Copilot for Pro+, Max, Business, and Enterprise users.
Use the Responses API for tool calling. Chat Completions works for this model only without tools. Do not send none or minimal effort, because neither is supported.
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-6.1-sol",
input="Find the retry-policy change in this diff and explain the blast radius.",
reasoning={"effort": "medium"}, # low, medium, high, xhigh, or max
)
print(response.output_text)
Final Thoughts
GPT-6.1 Sol delivers most of Astra's agentic performance at one-fifth of the token price, which makes it the right default for agentic coding, computer use, and business workflows.
OpenAI is betting that most developers do not need the frontier, only the frontier's output at a price they can run in a loop.
At $5.47 per Terminal-Bench Science 0.1 task against Astra's $23.80, that bet is stated plainly, and it puts pressure on Claude Opus 5.5 at $23.21. Anthropic's answer arrived a day earlier in Claude Sonnet 5.5, which lists at the same $2 and $10.
I'd move from GPT-6 Sol once you have handled the migration. The price is the same, and the exploit-development, research debugging, and low-effort factuality numbers are clearly better. But code that relies on none effort or on Chat Completions function calling has to change first.
I'd also keep a human review step on unattended agent runs, since coding deception (1.50% against Astra's 0.51%) and unwanted persistence (23.5% against 17.4%) both sit above Astra's rates.
I'd keep Astra for wet-lab reasoning, authorized security research, and the hardest scientific tasks, where the 15.5-point TroubleshootingBench gap and the 10-point ExploitBench Internal Port gap are real.
Nearly every figure here is OpenAI's own, with competitor results taken from public reports and no published independent reproduction. Run your own evaluation on your own workload before you commit.
If you want to build with models like this rather than just read about them, I recommend starting with our AI Fundamentals skill track.
FAQs
How does GPT-6.1 Sol compare to GPT-6 Astra?
How much does GPT-6.1 Sol cost?
Where can I access GPT-6.1 Sol?
What is the context window of GPT-6.1 Sol?
Is GPT-6.1 Sol safe to use in autonomous agents?
A senior editor in the AI and edtech space. Committed to exploring data and AI trends.
