Pular para o conteúdo principal

Claude Opus 5: Anthropic's New Flagship, Explained

An in-depth look at Claude Opus 5, including benchmarks, agentic performance, and how it compares to Fable 5, Opus 4.8, and rival models.
24 de jul. de 2026  · 7 min lido

Explorar com IA

Abrir no ChatGPTAbrir no ClaudeAbrir no Perplexity

Anthropic has a brand-new second-best model, again.

This time, it's Claude Opus 5, which is marketed as intelligence close to the frontier, but at half the price of Claude Fable 5 and, what's more, Opus 5 actually outperforms Fable 5 on several coding and knowledge-work evaluations, including Frontier-Bench and GDPval-AA. It's just second-place to the still not-widely-available Mythos 5. 

Below is a walkthrough of what's new with Opus 5, how it does on more of the benchmark tests, what Anthropic says about its alignment and safety profile, and what it costs.

What Is Claude Opus 5?

Opus 5 is the newest release in Anthropic's Opus tier, which sits, in terms of topline capability, below Mythos, but above Sonnet and Haiku.

Anthropic tells us its newest model is "thoughtful and proactive" (they always anthropomorphize a bit.) They also tell us the model is built for daily use, which to an AI engineer means it's efficient enough to run constantly.

On coding and knowledge-work evaluations such as Frontier-Bench and GDPval-AA, Anthropic says Opus 5 is now state-of-the-art among its models, though it still trails Mythos 5 specifically on cybersecurity tasks, which we cover in detail in our Mythos 5 article.

What's New with Claude Opus 5?

Two things stand out in Anthropic's own framing of what's actually new: a shift toward finishing tasks rather than stopping at a plausible-looking answer, and a jump in visual output quality.

Agentic persistence

Opus 5's biggest qualitative change is a willingness to keep working until a task actually succeeds. Anthropic gives us a picture of what that looks like:

  • Given a machine-part drawing but no way to view it directly, Opus 5 wrote its own computer-vision pipeline to extract the geometry from raw pixels.
  • Asked to fix a real bug in a popular open-source package manager, Opus 5 found and fixed the actual root cause, instead of a patch.
  • An engineer at a trading firm reportedly used Opus 5 to build a market-data feed for a new exchange in a single session — including writing its own test harness to validate the parsing logic.

This willingness to keep iterating rather than settle for a plausible-looking answer is what seems to be driving Opus 5's jump on agentic benchmarks like Frontier-Bench.

Stronger visual outputs

Anthropic also flags Opus 5 as capable of producing meaningfully stronger visual output than prior models, pointing to examples like a wind-tunnel-style visualization of airflow over aerodynamic objects, like a car. It was a compelling example, and good news if your job involves building generative diagrams.

Claude Opus 5 Benchmarks

This table only includes Anthropic's own models. I looked at their respective releases and lined up what I could. I'll compare Opus 5 with competing models in the next section.

A note: "not disclosed" or "not published" means that model wasn't included on that specific benchmark in the release it came from, not that it scored zero.

Benchmark Haiku 4.5 Sonnet 4.6 Opus 4.7 Opus 4.8 Sonnet 5 Fable 5 Opus 5
SWE-bench Verified 73.3% 79.6% 87.6% 88.6% 95.0% 96.0%
SWE-bench Pro 58.1% 64.3% 69.2% 63.2% 80.3% 79.2%
Terminal-Bench 41% 67.0% (2.1) 69.4% (2.0) 74.6% (2.1) 80.4% (2.1) 88.0% (2.1) not published
Frontier-Bench v0.1 (new eval) 18.7% 33.7% 43.3%
GDPval-AA / v2 (Elo, /2000) 1,753 1,890 (v1) / 1,593 (v2) 1,618 (v2) 1,747 (v2) 1,861 (v2)
Misaligned-behavior audit (lower = better) not disclosed not disclosed not disclosed 2.3 (best yet)
Price ($/MTok in/out) $1/$5 $3/$15 $5/$25 $5/$25 $2–3/$10–15 $10/$50 $5/$25

Claude Opus 5 vs. Fable 5

You can see in the table above that Opus 5 actually beats its own more expensive sibling on Frontier-Bench v0.1 (43.3% vs. 33.7%) and GDPval-AA v2 (1861 vs. 1747) - again, at half Fable 5's price.

Claude Opus 5 vs. Opus 4.8

This next one is the clearest generational jump: Opus 5 leads its predecessor on ARC-AGI-3, which jumps from 1.5% to 30.2%, and Frontier-Bench v0.1, which, as you can see on the table, above, more than doubles (43.3% vs. 18.7%). Keep in mind, Opus 4.8 was released only two months ago.

Claude Opus 5 vs. Competing Models

This next comparison was compiled by pulling benchmark results directly from the model release announcements published by Anthropic, OpenAI, and Google. I made a judgment to include rows where enough models had a published, comparable score to make the row meaningful. Also, benchmarks that were reported only as vague relative claims, I left out.

Category Benchmark Claude Opus 5 GPT-5.6 Sol GPT-5.6 Sol Ultra GPT-5.6 Terra GPT-5.6 Luna GPT-5.5 Gemini 3.6 Flash Gemini 3.5 Flash-Lite Gemini 3.5 Flash Gemini 3.1 Pro Preview
Knowledge Work GDPval-AA v2 (Elo) 1861 1736 1593 1591.8 1493.7 1421 1348.8 962.3
Coding SWE-Bench Pro 79.2% 64.6% 63.4% 62.7% 59.4% 54.2% 54.2%
Coding DeepSWE v1.1 68.8% 72.7% 69.6% 67.2% 67% 49% 37% (baseline) 11.8%
Coding Terminal-Bench 2.1 88.8% 91.9% 87.4% 84.7% 85.6% 54% (vs 3.1 Flash-Lite baseline 31%) 70.7%
Coding Frontier-Bench v0.1 (terminal coding) 43.3% 37.5%
Computer Use BrowseComp 90.8% 90.4% 92.2% 87.5% 83.3% 84.4% 85.9%
Computer Use OSWorld 2.0* 70.6% 62.6% 50.2% 45.6% 47.5%
Abstract Reasoning ARC-AGI-3 30.2% 7.78% 0.8% 0.18% 0.43% 0.42%
Tool Use AutomationBench 26.0% 18.1% 15.2% 14.9% 12.9% 14.5%

 

I should say: Every number here is self-reported by the lab that released the model, run on that lab's own evaluation harness, prompting setup, and (in some cases) effort or reasoning-level configuration.

Opus 5 vs. GPT-5.6 Sol

Opus 5 leads on the harder, more novel reasoning tasks (ARC-AGI-3, Frontier-Bench, AutomationBench) while GPT-5.6 Sol edges ahead on more established coding benchmarks like SWE-Bench Pro and DeepSWE. This might suggest Opus 5 may have an advantage in open-ended or less-templated problem solving, while Sol looks strong on well-defined, heavily-trained-for coding tasks.

Opus 5 vs. Gemini 3.6 Flash 

Gemini's recent release focuses on Flash-tier, cost-efficient models rather than a frontier flagship, so most rows either have no Gemini data or only Flash/Flash-Lite figures that trail both Opus 5 and GPT-5.6 by a wide margin. This comparison says more about the model tier than about Google's frontier capability.

Opus 5 vs. GPT-5.6

Where the two flagships overlap, Opus 5 leads on novel reasoning and agentic tasks (ARC-AGI-3, BrowseComp, OSWorld 2.0, AutomationBench), while GPT-5.6 Sol edges ahead on long-horizon coding (DeepSWE). Sol Ultra, Terra, and Luna mostly show up on separate coding benchmarks where Opus 5 has no published score, so this is really "Anthropic's one flagship vs. OpenAI's full family" rather than a clean single comparison.

Opus 5 Pricing and Availability

Opus 5 is available today across Claude's platforms at the same price as its predecessor:

  • $5 per million input tokens
  • $25 per million output tokens

Fast mode runs at roughly 2.5x the default speed, priced at twice Opus 5's base rate on the Claude Platform (or via usage credits in Claude Code) — the same structure as Opus 4.8.

Developers can access the model via the API using the model string claude-opus-5.

Two related features are launching in beta alongside Opus 5:

  1. Mid-conversation tool changes on the Claude Platform — developers can now swap which tools are available to Claude mid-conversation without invalidating the prompt cache.
  2. Automatic fallbacks — requests flagged by safety classifiers on Opus 5 or Fable 5 can now be routed automatically to another available model instead of being blocked outright, so API traffic keeps flowing to the best available model by default.

As with earlier Opus releases, Opus 5 carries no data retention requirements for general access.

Bottom Line on Claude Opus 5

If it feels like we're writing a lot recently about Claude models, it's because we are. Anthropic is heavily focused on making fast improvements to optimize model capability and efficiency. 

Now, Anthropic has produced something overall better than its previously best model, Fable 5. It comes with a meaningfully lower price, while also reporting its best alignment numbers to date. 

What's next? Well, the Haiku model is the oldest Anthropic model, and we're still waiting for an upgrade to the 5 series. That's one guess about what's next, but Anthropic keeps surprising us.


Josef Waples's photo
Author
Josef Waples

I'm a data science writer and editor with contributions to research articles in scientific journals. I'm especially interested in linear algebra, statistics, R, and the like. I also play a fair amount of chess! 

Tópicos

Learn with DataCamp

Curso

Introdução aos modelos Claude

3 h
12.3K
Aprenda a trabalhar com o Claude usando a API da Anthropic para resolver tarefas do mundo real e criar aplicativos com inteligência artificial.
Ver detalhesRight Arrow
Iniciar Curso
Ver maisRight Arrow
Relacionado

blog

Claude Opus 4.8: Anthropic's Smarter, More Honest Flagship Model

See what's new in Claude Opus 4.8: improved honesty, dynamic workflows, and effort controls.
Josef Waples's photo

Josef Waples

8 min

blog

Claude Opus 4.5: Benchmarks, Agents, Tools, and More

Discover Claude Opus 4.5 by Anthropic, its best model yet for coding, agents, and computer use. See benchmark results, new tools, and real-world tests.
Josef Waples's photo

Josef Waples

10 min

blog

Claude Opus 4.6: Features, Benchmarks, Hands-On Tests, and More

Anthropic’s latest model tops leaderboards in agentic coding and complex reasoning. Plus, it has a 1M context window.
Matt Crabtree's photo

Matt Crabtree

10 min

blog

Claude Opus 4.7: Anthropic’s New Best (Available) Model

Explore what's new in Anthropic's latest flagship: stronger agentic coding, sharper vision, and better memory across sessions. Compare the benchmarks against GPT-5.4, Gemini 3.1 Pro, and the locked-away Mythos Preview.
Josef Waples's photo

Josef Waples

9 min

blog

Claude Opus 4.8 vs GPT-5.5: Benchmarks, Tests, and Which to Choose

A head-to-head comparison of Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 across coding, reasoning, agentic tasks, and pricing.
Tom Farnschläder's photo

Tom Farnschläder

11 min

blog

Claude Opus 4.7 vs GPT-5.5: Which Frontier Model Is Best?

A head-to-head comparison of OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 across coding, reasoning, vision, tool use, and pricing.
Tom Farnschläder's photo

Tom Farnschläder

11 min

Ver MaisVer Mais