Google announced Gemini 4 Argon at the very end of September. Argon is Google's first Gemini 4 model and the new frontier. It arrives four weeks after OpenAI's GPT-6 Astra, OpenAI's new flagship since early September.
Google's launch table puts Argon ahead of Astra on most rows, at a fraction of Astra's price. There's a catch that matters more than any benchmark, though: Argon isn't generally available yet. Google is rolling it out first to vetted cybersecurity teams, with no date for developers or consumers.
TL;DR
- On Google's own benchmark table, Gemini 4 Argon beats GPT-6 Astra on most rows, led by legal, finance, and business workflow tasks.
- GPT-6 Astra keeps the edge on the hardest software engineering, scientific terminal work, and computer use.
- Argon's introductory price is a fifth of Astra's, and its standard price is still less than half.
- But you can't use Argon yet: access is limited to Google's trusted cyber defenders, with paid API customers and AI Ultra subscribers next.
- So, choose GPT-6 Astra if you need a frontier model in production today.
What Is Gemini 4 Argon?
Gemini 4 Argon is Google DeepMind's new frontier model, announced at the end of September. It's the first model in the Gemini 4 generation.
Google built Argon to sustain deep reasoning over long, multi-step work. That's backed by an output limit raised to 1M tokens so the model can reason through a hard problem in a single pass.
What Is GPT-6 Astra?
GPT-6 Astra is OpenAI's flagship model, released in early September. It sits at the top of the GPT-6 family above GPT-6.1 Sol and GPT-6 Luna.
OpenAI positions it as its most capable model for the most demanding work, from complex reasoning and coding to computer use and research.
Gemini 4 Argon vs GPT-6 Astra: Head-to-Head Comparison
Of the 19 benchmarks in Google's launch table, Argon scores higher than Astra on 14, Astra wins 4, and one is a tie. All of them come from Google's own table, so read them as Google's case for Argon rather than an independent ranking. I've pared that table down a bit to show what I think are the most well-known, interesting, and relevant benchmark tests for this release.
| Feature | Gemini 4 Argon | GPT-6 Astra |
|---|---|---|
| Vals Index | 68.9% | 63.1% |
| AutomationBench | 51.3% | 41.4% |
| Vals Finance Agent v2 | 65.4% | 53.5% |
| Harvey's Legal Agent Benchmark | 19.6% | 5.4% |
| DeepSWE v1.1 | 77.9% | 74.1% |
| FrontierSWE v2 | 55.0% | 65.5% |
| Terminal-bench 4.0 | 57.4% | 58.2% |
| Terminal-Bench Science 0.1 | 57.6% | 68.1% |
| OSWorld-2.0 (offline subset) | 69.2% | 72.6% |
| Max output tokens | 1M | 128K |
| Availability | Trusted cyber defenders only | Generally available |
Each dimension below carries the rest of the table.
Knowledge work: legal, finance, and business workflows
This is where Argon's lead is widest, and where Google aimed its launch.
On Harvey's Legal Agent Benchmark, which tests legal research and drafting, Argon scores 19.6% against Astra's 5.4%. That's a low score for both, but more than three times as many tasks completed is the biggest relative gap in the whole table.
The pattern holds across finance research and business automation, where Argon leads by roughly 10 to 12 points. If the claims hold up once people can use it, Argon is the stronger model for document-heavy legal and finance work.
Coding and software engineering
Coding is a split, and it depends on what kind of engineering you mean.
Argon leads DeepSWE v1.1, which tests long-horizon tasks in real codebases, and it edges Astra on Vibe Code Bench, which grades building working apps.

Astra wins clearly on FrontierSWE v2, and the two are within a point on Terminal-bench 4.0.
Google's own examples lean toward big migrations, including C and C++ to Rust across Google codebases.
For the hardest engineering benchmark in the table, though, Astra is still ahead.
Long context and visual understanding
Argon pulls away as inputs get longer.
On GraphWalks, which tests reasoning over graph structures spread across a long context, the two are close up to 128K tokens, but from 256K to 1M tokens, Argon scores higher.
Argon also leads on LVBench for long video understanding, while chart reading on Chartography is close to a tie.
For anyone feeding whole contracts, codebases, or recordings into a model, this is the most practical difference after availability.
Science, math, and computer use
Astra holds its ground on scientific terminal work, leading Terminal-Bench Science 0.1 by more than 10 points. Argon leads LABBench 2, a biology research benchmark, and RiemannBench in math. On computer use, Astra leads OSWorld-2.0 while Argon leads Agent's Last Exam, so neither owns this category.
Availability and safety
Astra is the one you can deploy today. Argon is currently limited to Google's Fairwind Program for trusted cyber defenders, plus Google's internal teams, and Google says paid API customers and Google AI Ultra subscribers come next without giving a date.
Argon's cyber story also cuts both ways. Google is releasing it to vetted defenders without cyber guardrails, and it ties Astra for first on CWE-bench v1, which tests fixing real security vulnerabilities.
For everyone else, Google says Argon will refuse harmful cyber and CBRN requests and runs monitors that can stop execution when the model oversteps the user's intent.
Pricing: what you actually pay
Argon costs a fifth of Astra at its introductory price and less than half at its standard price, but Google hasn't said how long the introductory price lasts.
Token rates side by side
| Rate | Gemini 4 Argon (introductory) | Gemini 4 Argon (standard) | GPT-6 Astra |
|---|---|---|---|
| Input, per 1M tokens | $2.00 | $4.00 | $10.00 |
| Output, per 1M tokens | $10.00 | $20.00 | $50.00 |
| Cached input read, per 1M tokens | $0.10 (95% off input) | Not published | $1.00 |
| Long-context surcharge | Not published | Not published | Over 272K input tokens: 2x input and cache, 1.5x output for the whole request |
The multiple is uniform on input and output: 0.2x Astra at introductory rates and 0.4x at standard rates. Both vendors bill reasoning as output, which matters more for Argon, since Google raised its output limit to 1M tokens precisely so it can think longer.
Even at standard rates, Argon comes in at well under half of Astra's cost. The open question is long prompts, where only Astra has published its pricing.
What a real workload costs
| Workload | Gemini 4 Argon (intro) | Gemini 4 Argon (standard) | GPT-6 Astra |
|---|---|---|---|
| Balanced assistant: 1M in / 250K out | $4.50 | $9.00 | $22.50 |
| Retrieval, each request under 272K: 10M in / 1M out | $30 | $60 | $150 |
| Retrieval, each request over 272K: 10M in / 1M out | Not published | Not published | $275 |
Each total is (volume ÷ 1M) × rate, summed across input and output, at standard rates.
- Balanced assistant: Argon saves $18 a month (80%) at introductory rates and $13.50 (60%) at standard rates.
- Retrieval under 272K: the same 80% and 60% savings, at a scale where they start to matter: $120 or $90 a month.
- Retrieval over 272K: Astra's surcharge pushes its cost to $275. Google hasn't said whether Argon has a long-context tier, so this is the row to check when Argon's pricing page goes live.
The two vendors use different tokenizers and neither has published token counts for a shared workload, so treat these totals as a rate comparison rather than a bill.
When to Choose Gemini 4 Argon vs GPT-6 Astra
For now, the fork is availability, not benchmarks: Astra is generally available, and Argon is limited to vetted cyber defenders. Once Argon opens up, its 0.2x to 0.4x price multiple will make it the default to test for knowledge work.
Choose Gemini 4 Argon if...
- Your work is legal research, financial analysis, or business automation. Those are its widest leads in Google's table, and the legal gap is the largest of all.
- You feed it very long inputs. It holds up far better than Astra on reasoning across 256K to 1M tokens, and leads on long video.
- Price matters at frontier quality. Even at standard rates, it costs well under half of Astra per token.
Choose GPT-6 Astra if...
- You need it today. Astra is in ChatGPT, the OpenAI API, Azure, Bedrock, GitHub Copilot, and OpenRouter, while Argon has no public API yet.
- Your hardest work is software engineering or scientific computing. It leads Argon on FrontierSWE v2 and Terminal-Bench Science by more than 10 points each.
- You run computer-use agents on desktop apps. It leads OSWorld-2.0, though Argon wins Agent's Last Exam, so test on your own tasks.
Final Thoughts
If you need a frontier model this week, GPT-6 Astra is the option. If your work is legal, finance, or long-document analysis, and you can wait, plan to test Gemini 4 Argon the day it opens up. We will keep you updated exactly when that is, if you sign up for The Median, our weekly newsletter:
This comparison, Argon versus Astra, is an interesting one. Google has launched Argon, along with a benchmark table showing that it beats OpenAI on most rows, with a price that undercuts Astra. But there's also no way for most people to use it.
Google is betting its lead holds until access arrives.
FAQs
Is Gemini 4 Argon better than GPT-6 Astra?
On Google's own benchmark table, Gemini 4 Argon scores higher than GPT-6 Astra on 14 of 19 benchmarks, with its biggest leads in legal, finance, business automation, and long-context work. GPT-6 Astra leads on FrontierSWE v2, Terminal-Bench Science 0.1, OSWorld-2.0, and Terminal-bench 4.0. All the scores come from Google, so independent tests are still needed.
Can I use Gemini 4 Argon yet?
Not unless you're in Google's Fairwind Program for trusted cyber defenders. Google says paid API customers and Google AI Ultra subscribers will get access next, but it hasn't given a date or published an API model ID.
How much does Gemini 4 Argon cost compared to GPT-6 Astra?
Gemini 4 Argon launches at an introductory $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward. GPT-6 Astra costs $10 and $50, so Argon is a fifth of Astra's price at introductory rates and 40% at standard rates. Google hasn't said how long the introductory period lasts.
Which is better for coding, Gemini 4 Argon or GPT-6 Astra?
It's split. Gemini 4 Argon leads DeepSWE v1.1 and Vibe Code Bench in Google's table, while GPT-6 Astra leads FrontierSWE v2 by about 10 points and edges Argon on Terminal-bench 4.0. For the hardest engineering tasks, Astra currently has the stronger published result.
What is the output limit of Gemini 4 Argon?
Google raised Gemini 4 Argon's output limit to 1M tokens, up from 64K on earlier Gemini models, so it can reason through long problems in a single pass. GPT-6 Astra's maximum output is 128K tokens.