Skip to content
Go To Agency
/AI & Tech
AI & Tech

Claude Opus 5: the best Opus yet, same price as Opus 4.8, approaching Fable 5, put to the test

On July 24, 2026, Anthropic shipped Claude Opus 5 at the exact same price as Opus 4.8 ($5 / $25 per million tokens) while getting close to Fable 5. New effort dial, strong company-reported benchmarks worth verifying, and our honest analysis of what these numbers are really worth.

By Robin MonteiroJuly 28, 202661 min · 13 434 mots
Claude Opus 5AnthropicAI modelsCoding agentsLLM benchmarks
Share article
Claude Opus 5: the best Opus yet, same price as Opus 4.8, approaching Fable 5, put to the test

It was on a Friday in late July, July 24 2026, that Anthropic shipped Claude Opus 5, live the same day across every one of its surfaces: Claude.ai, Claude Code, Claude Cowork, the API and the Claude Platform. Under the hood, one identifier to remember: claude-opus-5. The model slots in straight away as the default option on Claude Max and becomes the strongest selectable option on Claude Pro. Put another way, if you pay for an Anthropic subscription, there is a good chance your next conversations are already routing through Opus 5 without you touching a thing.

The official pitch fits in one sentence, which we quote verbatim: Opus 5 approaches the frontier intelligence of Claude Fable 5 for half the price. It is a clever line, and it deserves a careful read. Anthropic does not claim Opus 5 is its smartest model: that title stays explicitly with Fable 5, the flagship. Opus 5 is positioned as the daily driver, the everyday workhorse with the best cost-to-performance ratio, the one you reach for by default on complex agentic coding and enterprise work before you roll out the heavy artillery. The message carries a whiff of marketing, and the word approaches does all the heavy lifting: approaching the frontier is not reaching it, and "half the price" is worth a close look at what you actually get for those savings.

This is exactly the kind of claim we do not take at face value. To keep this article at arm's length from the press release, we sort every piece of information into one of three buckets, never blurring the lines. One: the verified facts, what is objectively observable (model ID, context window, list price, availability dates). Two: the benchmarks reported by Anthropic, those impressive numbers pulled from the announcement and the system card, which are interesting but remain results communicated by the vendor itself, often on its own runs, so to be taken with the usual caution. Three: our analysis, what we at Go To Agency make of it after cross-checking third-party sources and getting our hands on the model. When a figure comes from Anthropic, we say so; we will never present it as an independent truth.

Three new features structure this launch, and we will detail each in its own section. First, price: Opus 5 costs exactly the same as Opus 4.8, that is $5 per million input tokens and $25 per million output tokens. Anthropic froze the tariff and raised the capacity, which, in a price war against OpenAI and Google, says a great deal about the strategy. Next, the effort dial (the high, xhigh, max levels), a slider that lets you explicitly trade off cost against capability, request by request. Finally, automatic fallback: when the model declines a request on safety grounds, the API switches on its own to another model rather than handing you back an error. Three signals that, taken together, sketch a model built for production and for budgets, not just for lab records.

Before we open the hood on these mechanisms, let us lay the foundations. The official Claude Platform documentation provides the model's exact specifications: context window, maximum output, input modalities, knowledge cutoff date. Let us start with the verified spec sheet.

Anthropic Claude Opus 5, a luminous structured intelligence form representing the new default agentic model, launched July 24, 2026

The verified spec sheet

Before the analysis, the raw facts. Here are the official specs for claude-opus-5 as Anthropic documents them at the July 24 2026 launch. Not a single figure in this table is a benchmark or a marketing projection: these are the verifiable technical specs, the ones you will find in the official Claude platform documentation.

SpecificationDetail
API identifierclaude-opus-5 (pinned snapshot, not an evergreen pointer)
Cloud availabilityClaude API, AWS Bedrock (anthropic.claude-opus-5), Google Cloud Vertex AI (claude-opus-5), Microsoft Foundry
Release dateFriday July 24 2026 (available immediately)
Context window1 million tokens (roughly 555,000 words, or about 2.5 million unicode characters)
Maximum output128,000 tokens (up to 300,000 via the Batch API, beta header output-300k-2026-03-24)
Knowledge cutoffMay 2026 (the most recent in the Claude 5 lineup)
Training data cutoffMay 2026
ModalitiesText and image input (vision), multilingual. Text output only
Adaptive thinkingYes (reasoning active by default)
Extended thinking (thinking.type enabled)No
Comparative latencyModerate
Default efforthigh on the Claude API and Claude Code

Two rows in this table are worth pausing on, because they concretely change what you can do with the model. The first: the May 2026 knowledge cutoff. It is the freshest knowledge in the entire Claude 5 lineup, ahead of Sonnet 5, Opus 4.8 and Fable 5, all fixed on January 2026. Four months of difference may sound trivial, but for a model built for agentic coding and enterprise work, it matters: software libraries, framework versions, regulatory changes, the news cycle of a given sector. An agent that has to reason about a recent technical ecosystem starts with fewer blind spots, and therefore fewer hallucinations to correct. This is not a cosmetic detail, it is a reduction in verification debt.

The second: the one million token context window. Roughly 555,000 words, the equivalent of several large code repositories, a complete client file or hundreds of pages of documentation ingested in a single request. For an agent working on long tasks (crawling a monorepo, holding a conversation thread that runs for hours, keeping an entire spec in mind), this working memory is the number one limiting factor. Output, for its part, is capped at 128,000 tokens in standard use, which stays generous for code or a report, and climbs to 300,000 via the Batch API for massive generations.

The pairing of adaptive thinking set to Yes and extended thinking set to No reflects an architectural choice. Opus 5 reasons by default: its thinking mechanism is active and tunable through the effort parameter (which we detail further down), without having to switch on a separate extended reasoning mode as on previous generations. Reasoning is no longer an option you plug in, it is the model's native behavior. For the integrator, that simplifies configuration: you set the effort dial, not a reasoning toggle.

Last point, technical but decisive for anyone putting this model into production: the claude-opus-5 identifier is a pinned snapshot, not an evergreen alias that would silently drift toward a newer version. Since the 4.6 generation, Anthropic has adopted this identifier format with no visible date but pinned to a precise version. In plain terms: the model behind this ID will not shift under your feet. That is what guarantees the reproducibility of your evaluations, the stability of your prompts and the absence of surprise regressions in the middle of a production run. For an agency like ours, which delivers integrations to clients, this stability is not a comfort, it is a condition of being taken seriously: you test a version, you validate it, you deploy it, and it stays identical until we decide to migrate ourselves.

Pricing: identical to Opus 4.8, half of Fable 5

Let's start with the number that anchors everything else. Claude Opus 5 is priced at $5 per million input tokens and $25 per million output tokens. Those are exactly the same rates as Opus 4.8, its predecessor. Anthropic left the pricing untouched: same cost as before, for a meaningfully higher capability (the detailed benchmarks come later in this article). In a sector where every new model generation usually arrives with a price premium, holding the line while moving up the ladder is a choice worth flagging.

The other useful reference point: this price is exactly half that of Claude Fable 5, the most capable widely available model in Anthropic's lineup, billed at $10 input and $50 output per million tokens. That is the heart of the official pitch. Anthropic frames Opus 5 as a model that, in its own words, approaches Fable 5's frontier intelligence at half the price. Note the careful wording: the company does not claim Opus 5 is its smartest model (that title stays with Fable 5), only that it comes close to those numbers at a lower cost. This is a daily-driver positioning, the everyday model for agentic work and enterprise, with Fable 5 reserved for cases that demand maximum capability.

Output price per million tokens, current Claude lineup (lower = cheaper)

Haiku 4.5
$5
Sonnet 5
$15
Opus 5
$25
Opus 4.8 (legacy)
$25
Fable 5
$50

Source: Anthropic model and pricing documentation, July 2026 (verified). Output price per million tokens.

The chart puts Opus 5 back in the context of the full lineup. At the very bottom, Haiku 4.5 remains the budget workhorse ($1 input, $5 output), built for volume and simple subtasks. Sonnet 5 sits in the middle at $3 / $15 (with an introductory rate of $2 / $10 through August 31 2026). Opus 5 and Opus 4.8 share the same rung at $5 / $25, and Fable 5 doubles the stakes at $10 / $50. The read is clear: at near-frontier capability, Opus 5 halves the output bill compared to Fable 5, and that is precisely where the trade-off gets decided for most real-world workloads.

Fast Mode: twice the price, two and a half times the speed

Anthropic offers a speed option called Fast Mode. In practice, it serves the same model at roughly 2.5 times the default speed, but bills 2 times the base rate, that is $10 input and $50 output per million tokens (curiously, the exact standard price of Fable 5). The math is simple to lay out: you pay a latency premium. For an interactive user experience (a live coding assistant, a chatbot where every second of waiting counts), spending double to deliver the answer two and a half times faster can be justified. For background batch processing, where latency does not matter at all, it would be money thrown out the window. Fast Mode is a product-experience lever, not a cost-savings lever.

Lower the effort, lower the bill

Opus 5's real cost lever lies elsewhere: it is the effort dial. We cover how it works technically further down, but its impact on cost deserves to be raised at the pricing stage, because it changes how you reason about spend. The effort parameter (levels low, medium, high, xhigh, max) sets how much compute and how many reasoning tokens the model burns on a given task. The economic logic is direct: lowering the effort cuts token consumption, and therefore the bill, while preserving most of the performance on tasks that do not demand the deepest reasoning.

Concretely, the price per million tokens does not move, but the number of tokens spent solving a problem drops when you step down a notch. Triaging tickets, rephrasing a message, classifying documents: none of these need the same thinking budget as a multi-file refactor. Where Opus 4.8 forced high effort everywhere, Opus 5 hands the trade-off back to you: by default it settles on high on the API and in Claude Code, but nothing stops you from dropping to medium or low on low-stakes tasks and reserving xhigh or max for the problems that warrant it. The real cost of a workflow no longer reads off the pricing grid alone: it is steered at the request level.

Why holding Opus 4.8's price is aggressive

Our analysis. Freezing the price from one generation to the next is anything but trivial in July 2026. Anthropic is launching Opus 5 into a market in the middle of a price war, up against OpenAI's GPT-5.6 "Sol" and Google's Gemini 3.1 Pro, both pushing their capability-per-dollar, not to mention the Chinese labs slashing rates at the entry level. In that context, keeping $5 / $25 while raising capability appreciably amounts to lowering the cost per unit of useful work without touching the sticker price. It is a disguised price cut, and a head-on response to competitive pressure.

The timing reinforces the read. Anthropic is preparing a public offering later this year, and a cheap "daily driver" model that captures enterprise workloads at scale is exactly the kind of product that strengthens a growth story ahead of an IPO. As Anthropic's product lead for research puts it, what enterprises want is value: a cheaper model that does not reach comparable quality is not actually useful. Holding Opus 4.8's price with higher capability plays both hands at once.

The caveat to keep in mind: the tokenizer

Our analysis. A sticker price is not a bill. The Opus 4.7 generation and later (including Opus 5) use a tokenizer that produces, for the same text, roughly 30% more tokens than models predating 4.7. In other words, the same prompt and the same response count as more tokens than they would with an older generation, which mechanically inflates the effective cost at constant content. This is not a trap, just a billing reality to factor in: if you are migrating from a pre-4.7 model, do not compare prices per million tokens assuming an identical token volume. The only reliable way to know your real cost is to measure it on your own workload, with your typical prompts and outputs, then compare end-to-end bills rather than pricing grids. That is the nuance separating an informed purchasing decision from a nasty surprise at month's end.

The real cost of running an agent on Opus 5

On paper, Opus 5 keeps the same pricing grid as Opus 4.8: $5 per million input tokens, $25 per million output. But for an agency running a coding agent around the clock, it is the real-world volume that lands on the invoice. Take a concrete, realistic scenario: an agent that processes roughly 2 million input tokens per day (context, files, history) and 400,000 output tokens (patches, explanations, commands).

Three models, same workload

The math is simple. Input: 2 x $5 = $10. Output: 0.4 x $25 = $10. That comes to $20 a day, roughly $600 a month for this agent on Opus 5. Here is how the three models relevant to this kind of task compare.

ModelInput costOutput costTotal per dayTotal per month (approx)
Opus 5$10$10$20$600
Fable 5$20$20$40$1,200
Sonnet 5$6$6$12$360

For the same workload, Fable 5 costs double what Opus 5 does, and Sonnet 5 comes in about 40% cheaper. The monthly gap between the extremes tops $800 on a single agent: multiply that across a team of developers and the choice of model becomes a budget line in its own right.

The effort dial shifts the bill

Opus 5 exposes an effort dial (levels high, xhigh, max, with high as the default on the API and in Claude Code). Dialing the effort down mainly cuts the number of output tokens, and output is the most expensive line at $25 per million. In practice, dropping from high to a moderate effort on simple tasks can shave off a meaningful chunk of the $10 in daily output spend without touching the input side.

Conversely, the max lever is for the hard problems. Anthropic reports that on its own CursorBench 3.2 (an in-house benchmark, not an independent measure), Opus 5 at max effort lands within 0.5% of Fable 5, for roughly half the cost per task. In other words, max effort often lets you skip paying for Fable 5 while keeping near-identical performance.

Three concrete levers for optimization

  • Measure your actual tokens. The Opus 4.7 generation and later produces roughly 30% more tokens than models before 4.7, because of the tokenizer. Do not trust a theoretical estimate: instrument your own workload before you budget.
  • Route volume to the smaller models. Send the bulk of your traffic (classification, summaries, routine tasks) to Sonnet 5 ($3 / $15, intro pricing at $2 / $10 through August 31 2026) or Haiku 4.5 ($1 / $5), and keep Opus 5 or Fable 5 for the genuinely hard work.
  • Lean on Batch and prompt caching. The Batch API (output up to 300k tokens via a beta header) is often significantly cheaper for non-urgent jobs, and prompt caching cuts input cost whenever the same context is reused from one call to the next.

The effort dial: the real headline feature

If you take away one product change from this launch, make it this one. Since Opus 5, you no longer just choose which model to call, you choose how hard you want it to think. In the API this setting has a name: the effort parameter, passed through output_config.effort. It is the central lever of the model's economics, and it is also what explains how Anthropic manages to sell intelligence close to its frontier model for half the price.

In concrete terms, effort arbitrates between two resources: the intelligence applied to a task and the number of tokens (and therefore the cost and latency) the model burns to get there. The higher the effort, the longer Opus 5 thinks, the more paths it explores, the more it checks its own work. The lower it is, the faster and cheaper it answers while preserving, and this is the key point, most of its performance. This is not a plain quality slider that degrades the output in a straight line: dropping one notch cuts the cost far faster than the performance.

Three levels per the vendor, five in the docs

Watch out for a bit of confusion doing the rounds in the press. Fortune, in its launch story, describes a simple toggle between cost and capability and mentions three levels (low, medium, high). Anthropic's own announcement foregrounds three step-up tiers: high, xhigh and max. The technical documentation, however, exposes a full five-notch scale: low, medium, high, xhigh, max. When in doubt, the docs win: the Fortune piece is a mass-market simplification.

Two behaviors worth knowing. First, the default is high on the API as well as in Claude Code: explicitly passing effort: "high" is exactly the same as passing nothing at all. That is a quiet but real change from Opus 4.8, which forced high everywhere, including on claude.ai. Second, on Opus 5 the setting drives the volume of thinking, not the length of the visible answer: if you want a short reply, the thing to lower is not effort, it is your prompt, which needs to ask for it.

Why this lever changes everything for an agentic budget

The effort dial does not carry the same weight for a one-off chatbot as it does for a coding agent. An agent that runs for thirty minutes on a repo, reading files, launching tests, fixing and starting over, burns tokens by the million. At that rate, a factor of two or three on consumption is no accounting footnote: it is the difference between a profitable automation and an end-of-month bill that blows up. That is exactly the problem the dial is there to address.

The clearest illustration comes from CursorBench 3.2. At max effort, Opus 5 sits, a figure declared by Anthropic, within 0.5% of Fable 5's top score, but for half the cost per task. In other words, the dial lets you reach near-frontier performance on the tasks that deserve it, without paying Fable 5 rates on every call. The same logic holds in reverse: on simple sub-tasks or high-volume agents, dropping to low or medium frees up most of the budget for the moments that truly matter.

When to turn it up, when to stay low

The practical rule comes down to a handful of cases. Go to max when you are betting big on a single shot: a genuinely hard piece of reasoning, a critical one-shot migration, a research problem or an analysis you do not want to regenerate. That is the notch where the model spends without a token budget to maximize its chances on the first try. Move up to xhigh for long-haul agentic work: tasks over thirty minutes, budgets in the millions of tokens, autonomous coding on sprawling codebases.

Stay at high for the bulk of day-to-day work: it is already a serious level of capability, and it is the default for good reason. Drop to medium or low for sub-agents, bulk triage, and steps where speed and cost outrank finesse. One piece of advice worth its weight in gold, and one the documentation hammers home: do not recycle the settings you inherited from your old models. Run a fresh effort sweep against your own evaluations, because the right balance point depends entirely on your actual workload.

Switching models mid-task

The effort dial has a product-side cousin: the ability to trade cost against capability no longer inside a single model, but between models, over the course of one task. That is the toggle Fortune sums up with its cost-versus-capability formula. In practice, a well-designed orchestration runs a powerful model on the planning or heavy-reasoning steps, then hands off to a cheaper one (Opus 5 itself against Fable 5, or a notch below) for high-volume implementation or review. You only pay the premium rate where it produces value.

The overall lesson is this: with Opus 5, cost is no longer a fixed property of the model you call, it is a variable you steer call by call, notch by notch. For an agency or a product team industrializing its AI usage, this is arguably the most concrete change of the launch: the same technical building block can cost three times less depending on how you tune it, without ever switching models.

The Claude Opus 5 effort dial, a control balancing cost against capability across the high, xhigh and max levels

Automatic fallback and tools you can swap mid-conversation

Beyond the effort dial, Opus 5 ships with two product features that speak mainly to teams running agents in production. They are less flashy than a benchmark score, but they concretely change the reliability and the cost of what you deploy. Let us look at them one at a time, through the eyes of the developer who has to own the uptime.

Automatic fallback: an answer instead of an error

Opus 5 (like Fable 5) carries safety classifiers that can refuse a request. Key thing to understand so you do not get your code wrong: a refusal is not an error. On the API side, it comes back as an HTTP 200 with stop_reason: "refusal" and a stop_details object that spells out the category and an explanation. In other words, your try/catch will not catch it: you have to inspect the response stop_reason.

The new part is the fallback behaviour. As Fortune describes, when a request is refused the API can automatically switch to another model so the user gets an answer rather than an error. On Claude.ai, Claude Code and Claude Cowork, this fallback goes to claude-opus-4-8 by default. On the API, you have to enable it explicitly. The simplest mode is called "default": you set fallbacks: "default", and for each refusal category (the categories are cyber, bio, frontier_llm, reasoning_extraction, general_harms) the API replays the request on the model Anthropic recommends. You no longer have to maintain your own list of backup models. One concrete detail for biology: requests that used to be blocked on Fable 5 are now routed to Opus 5.

Technically, the response traces the handoff cleanly: the top-level model field names the model that actually served the request, a content block of type fallback marks the switch (from which model to which model), and usage.iterations keeps the breakdown of each attempt along with its billing. Worth knowing too: after a fallback, the API pins that conversation prefix to the model that accepted for roughly an hour, so it does not re-attempt a predictable refusal on every turn.

The upside. In production, this is pure robustness. A generation funnel, a support agent, a document-processing pipeline: none of these flows should stop dead because a classifier raised a false positive on a perfectly legitimate request. Fallback turns a potential break into graceful degradation. The end user does not see an error, they see an answer, produced by a slightly less capable but functional model.

The downside. This comfort has a cost that is not monetary: determinism. Your system can, with no warning on the caller side, serve an answer from a model other than the one you asked for. If you do not systematically log the response model field and the contents of usage.iterations, you lose track of who answered what. In practice, our recommendation to teams: always log the serving model, alert on an abnormal fallback rate (it is often the signal that a prompt is grazing a safety boundary), and test your critical paths with the fallback model in mind, not just with Opus 5. One more point of vigilance: billing. Each attempt is billed at its own model's rate, which makes the cost of a refused-then-fallen-back request less predictable than a plain call.

Swapping tools without breaking the prompt cache

The second feature is quieter but, for anyone building long-running agents, potentially the more profitable of the two. On the Claude Platform, you can now add or remove tools in the middle of a conversation without invalidating the prompt cache. It is a beta feature, turned on by a dedicated header.

To grasp what is at stake, you have to remember how the cache works. Tool definitions are sent at the head of the prompt. Until now, the slightest change to that tool list changed the prompt prefix and therefore invalidated the cache: the model had to re-read the entire context at full price, instead of benefiting from the reduced rate of a cache read. For an agent that runs thirty minutes or more, with a context that keeps growing and a toolset that shifts by stage (research first, then file writing, then execution), that permanent invalidation was a sink for tokens and latency.

With hot tool swapping, the prefix stays stable and the cache stays warm even when the toolbox changes. The gain is twofold: fewer billed tokens (a cache read costs a fraction of a rewrite) and less latency on every turn, since the model does not re-pay the cost of ingesting the whole history. For a developer orchestrating a multi-step agent, this is the difference between exposing a huge tool catalogue up front (which pollutes the context and raises the risk of a bad call) and serving exactly the relevant tools at each phase, with no cache penalty. Anthropic has also lowered the minimum cacheable prefix to 512 tokens on Opus 5 (versus 1,024 on Opus 4.8), which widens the field of short prompts that become cacheable.

Our read. Taken together, these two features tell the same story: Anthropic is tuning Opus 5 for real, durable agentic use, not just for the benchmark score. Fallback targets resilience in production, hot tool swapping targets the cost and latency of long loops. These are exactly the two pain points every team hits when it moves from an agent demo to an agent that runs for real, for real customers.

Claude Opus 5 benchmarks, performance bars comparing SWE-bench Pro against Opus 4.8, Fable 5 and Mythos 5

The benchmarks Anthropic reports

Keep one reflex handy as you read what follows: every number in this section comes from Anthropic. They are drawn from the official announcement and the system card published on July 24 2026. These are company-reported results, meaning they were measured and communicated by the model's own maker, on configurations and test sets it selected itself. None of these runs was conducted by an independent third-party evaluator. They point to a direction, not a truth carved in stone. We report them because they sketch the positioning claimed for claude-opus-5, but we treat them with the same caution as a spec sheet signed by the manufacturer.

The through-line of the announcement is simple: Opus 5 aims to approach the intelligence of Fable 5 for roughly half the cost per task. Almost every benchmark put forward is therefore built around the same argument, performance relative to price, not raw performance. It is a telling editorial choice: Anthropic does not claim Opus 5 is its smartest model (that title still belongs to Fable 5), but that it delivers the best return on every euro spent. Here are the reported results, presented as they appear in the maker's own communication.

Benchmark (reported by Anthropic) What it measures Reported Opus 5 result
Frontier-Bench v0.1 Agentic coding on long-horizon tasks Beats every other model and doubles Opus 4.8's performance, at a lower cost per task
CursorBench 3.2 (effort max) Coding inside an IDE, editing real code Within 0.5% of Fable 5's best score, for half the cost per task
ARC-AGI 3 Abstract reasoning, generalization Scores about 3 times higher than the next best model
Zapier AutomationBench End-to-end business workflow automation About 1.5 times the success rate at equal cost (Zapier says it reached the top of its leaderboard)
OSWorld 2.0 Computer-use, driving a real computer Beats Fable 5 for a third of the cost
Organic chemistry (internal eval) Scientific reasoning, chemistry +10.2 percentage points vs the predecessor
Protein tasks (internal eval) Biology, protein modeling +7.7 percentage points vs the predecessor

Read together, these benchmarks tell a coherent story. Frontier-Bench v0.1 and CursorBench 3.2 measure agentic coding, that is, the model's ability to chain actions inside a code repository, edit multiple files and carry an engineering task over time, rather than simply completing a snippet of a function. This is the arena Anthropic absolutely wants to win, and the message is twofold: Opus 5 would do twice as well as the previous generation on Frontier-Bench, and come within a hair of Fable 5 on CursorBench for half the price. It is worth flagging that the best CursorBench result is obtained at effort max: that is the most token-hungry configuration, and therefore the most expensive in real usage, even if the cost per task is still presented as favorable.

ARC-AGI 3 shifts register. This family of tests targets abstract reasoning and generalization, the ability to solve novel problems the model could not have memorized during training. A score announced as three times higher than the nearest competitor is spectacular on paper, but it is also the kind of figure that calls for the most caution: the ARC-AGI leaderboard moves fast, version 3 is recent, and a gap of this size still needs to be confirmed by third parties before any firm conclusion can be drawn.

Zapier AutomationBench and OSWorld 2.0 follow a different logic, that of applied agentics and computer-use: having the model run complete business workflows, or driving an interface, clicking, filling in fields and navigating the way a human would in front of a screen. The success rate of about 1.5 times higher at equal cost on AutomationBench, backed by Zapier's testimonial saying it reached the top of its leaderboard, and beating Fable 5 for a third of the cost on OSWorld 2.0, both serve the same commercial argument: Opus 5 would be the model you plug into real automations without blowing your token budget.

Finally, the last two rows step outside coding entirely to touch scientific research. The gain of 10.2 percentage points in organic chemistry and 7.7 points on protein tasks, measured internally against the predecessor, feed Anthropic's claim that Opus 5 is its most capable generally available model for science, with a particular strength in biology. The wording matters: these evals are internal, the baseline for comparison is Anthropic's own previous model rather than a rival, and the gap is expressed in raw points without the protocol details being made public.

The underlying limitation is the same across the whole list: these are metrics chosen by the maker, on benchmarks whose selection, effort configuration and measurement mode it decides. Nothing dishonest in itself, it is the norm for a launch, but it remains a self-assessment exercise. The real arbiter will come from independent runs and field feedback, which we examine in the following sections. For the full breakdown of the numbers, the official Claude Opus 5 announcement remains the reference source to consult directly.

Third-party benchmarks: what the numbers are actually worth

A word of caution before any figure: the benchmarks cited here are genuinely built by organizations outside Anthropic (academic researchers, specialized evaluators, code-review platforms), which makes them more credible than in-house tests. But there is an under-discussed reality worth flagging at launch: the runs that produced the July 24 2026 scores were carried out by Anthropic on those test sets, not by third parties in a blind setup. The benchmark is independent; the way it was run is not entirely. That is standard practice for a launch (the vendor wants numbers to communicate on day one), but it remains a limitation: a score reproduced three months later by a neutral lab, on its own harness and with its own prompts, will not necessarily land in the same place.

SWE-bench Verified: 96.0%, the showcase number

The most widely quoted figure is the one from SWE-bench Verified, the de facto reference for measuring how well a model resolves real GitHub tickets (actual issues fixed by humans, with validation tests). On that bench, claude-opus-5 hits 96.0%, a number reported by BenchLM as the average of five runs. At this level, we are talking about saturation: the benchmark barely discriminates between the top models anymore, since only a handful of tickets are left to fail on. A 96% score is impressive, but mostly it tells you that SWE-bench Verified is reaching the end of its life as a measuring instrument. That is exactly why serious evaluators have shifted to harder variants.

SWE-bench Pro: 79.2%, the third step on the podium

This is where the reading gets interesting. SWE-bench Pro keeps the real-ticket principle but on markedly nastier problems, designed to resist memorization and overfitting. On it, Opus 5 lands 79.2%, a figure confirmed identically by BenchLM and by Codersera. The point to hold onto is not the raw score, it is the ranking: Opus 5 comes in third overall, behind Mythos 5 (80.3%) and Fable 5 (80.0%), but well above Opus 4.8, which topped out at 69.2%.

Translation: the generational gain is real and substantial (ten points better than its direct predecessor on a hard bench), but there is no magic leap that would put Opus 5 ahead of Anthropic's high end. On this test the daily-driver model stays a hair from the frontier (eight tenths of a point separate Opus 5 from Fable 5) yet clearly behind once you look at the actual step. The marketing promise (frontier performance at half the price) holds on the cost-to-capability ratio, not on absolute capability. The chart below makes that ranking visible.

SWE-bench Pro, reported score (higher is better)

Opus 4.8
69.2%
Opus 5
79.2%
Fable 5
80.0%
Mythos 5
80.3%

Source: SWE-bench Pro scores reported at launch (July 2026). Runs carried out by Anthropic, not fully independent.

Senior SWE-bench: where Opus 5 truly dominates

The most revealing bench for professional use is not the most publicized one. On Senior SWE-bench, the analysis by Snorkel places Opus 5 at the top of every model evaluated in the bug investigation and performance category. This is the kind of task a senior engineer does day to day: understanding why a system misbehaves, tracing a chain of causes, isolating the true root rather than the symptom. That the model comes out first on precisely this register is more meaningful, for an agency or a product team, than one extra point on a saturated bench. It is not a promise of full autonomy, but a signal that the model is strong where the work is expensive.

Snorkel: the fault lies in the inference, not the formatting

The same Snorkel analysis surfaces what raw scores always hide: why the model fails when it fails. By dissecting the failed trajectories and confirming root causes, Snorkel establishes that faulty inference accounts for 35% of failures, by far the leading mechanism. In other words, when Opus 5 gets it wrong, it is almost never because it mis-formatted its output or mishandled a tool (those categories stay marginal), it is because it reasoned poorly: it drew a false conclusion from elements that were actually available. That is actionable information in production. It means a formatting guardrail or a more robust parser will not rescue much; what protects you is human verification of the reasoning on high-stakes tasks, and breaking complex problems into steps where a bad deduction becomes easier to spot.

CodeRabbit: the precision versus recall trade-off laid bare

The most instructive test comes from CodeRabbit, which evaluates models on concrete ground: automated code review. Their verdict on Opus 5 in x-high configuration cuts both ways, and that is exactly what makes it credible. On the plus side, the model produces a more precise stream of comments: 39.3% of its comments are genuinely actionable and relevant, versus 35.2% for CodeRabbit's production baseline. Put plainly, when Opus 5 flags something, it is right more often.

The flip side is stark: on that same test, Opus 5 catches fewer known issues from the benchmark, with 55.2% coverage against 61.1% for the baseline. It is more accurate, but it lets more real defects slip through. This is the classic precision-versus-recall trade-off: a reviewer who speaks up less but more accurately, versus a reviewer who casts a wider net at the cost of more noise. Neither stance is intrinsically better; it all depends on your pipeline. On a repo where every false positive costs human time, Opus 5's precision is an asset. On a security audit where missing a defect is expensive, the baseline's superior coverage matters more. The meta-lesson lies elsewhere: a single score never tells you whether a model suits you, you have to look at the error profile.

What to take away

Three honest observations emerge from these third-party evaluations. First, the progression over Opus 4.8 is undeniable and wide: ten points on SWE-bench Pro, top of the ranking on bug investigation, that is not marketing. Second, there is no leap above the frontier: on the hard benches, Opus 5 stays behind Fable 5 and Mythos 5, which is consistent with Anthropic's stated positioning (best cost-performance, not the smartest model). Finally, the CodeRabbit detail is a reminder that a gain in precision can be paid for in recall: the aggregate figure flatters, the error profile informs. For anyone who has to choose a model in production, it is these three nuances, and not the showcase 96.0%, that should guide the decision.

Opus 5, Fable 5, Mythos 5, Opus 4.8: who does what

With four models carrying the number 5 and an Opus 4.8 still lingering in the docs, Anthropic's lineup has become hard to read at a glance. The good news: the logic is actually simple once you lay it out flat. Each model sits in a precise slot along two axes, cost and capability, and Opus 5 reshuffles the deck not at the top of the stack but in the middle. Here is the lineup as it stands at the end of July 2026.

Model Price in / out (per million tokens) Context Knowledge cutoff Positioning
claude-haiku-4-5 $1 / $5 200k February 2025 Speed and volume, simple tasks
claude-sonnet-5 $3 / $15 (intro $2 / $10 through August 31 2026) 1M January 2026 Balance of speed / price / quality
claude-opus-4-8 (legacy) $5 / $25 1M January 2026 Former daily driver, replaced by Opus 5
claude-opus-5 $5 / $25 1M May 2026 Daily driver, best cost-performance
claude-fable-5 $10 / $50 1M January 2026 Maximum capability, the smartest
claude-mythos-5 $10 / $50 1M January 2026 Defensive cyber, invite-only (Project Glasswing)

The lineup logic, read top to bottom

Anthropic makes no secret of its hierarchy. Fable 5 remains the smartest model in the house, the one you reach for when you need maximum capability, and Opus 5 makes no claim to that title. The vendor's official line is explicit: Opus 5 is "a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price", in other words a model that approaches Fable 5's frontier intelligence for half the cost. The key word is approaches: it does not match it, it gets close enough that, on most real tasks, the gap no longer justifies doubling the bill.

Below it, Sonnet 5 plays the balancing act: cheaper ($3 / $15, with an introductory rate of $2 / $10 through August 31 2026) and faster, it absorbs the bulk of application volume when the difficulty stays moderate. And at the very bottom of the stack, Haiku 4.5 leans into throughput and floor cost ($1 / $5) for simple tasks, classification, high-volume sub-agents, and anything that needs to answer in near real time. It is the only one in the lineup still stuck on a 200k context and a February 2025 knowledge cutoff, which gives away its relative age.

Mythos 5 is the odd one out. Same specs and same price as Fable 5 ($10 / $50), but it is not generally available: access is by invitation only, through Project Glasswing, a program geared toward cyber defense and critical infrastructure. For nearly every reader of this blog, Mythos 5 is not an option you choose: it appears in the lineup only as a capability reference point. We leave the detail of its security posture to another section of this article; here it is enough to note that it occupies a slot you will, barring exceptions, never have access to.

The real message: Opus 5 replaces Opus 4.8, not Fable 5

This is the most misunderstood point of the launch, and the one that changes everything in an architecture decision. Opus 5 does not move up a tier, it replaces Opus 4.8 like for like on the billing side. Same price ($5 / $25), same context window (1M tokens), same Opus family. What moves is the performance inside that pricing envelope: Anthropic held the price and raised the capabilities, with a more recent knowledge cutoff thrown in (May 2026 versus January 2026 for Opus 4.8). Opus 4.8 is not retired for all that; it drops to legacy status and stays available, notably as the target of the automatic fallback described earlier in this article.

In practical terms, if you were running Opus 4.8 in production, migrating to Opus 5 is an upgrade with no extra cost per token: you pay the same, you get better, and you get a better-informed model. Anthropic's documentation even frames it as a recommendation: start with Opus 5 for complex agentic coding and enterprise work, and move up to Fable 5 only when you need absolute maximum capability. Opus 5 has, moreover, become the default model on Claude Max and the strongest option selectable on Claude Pro, which confirms its calling as the everyday workhorse.

The dividing line between Opus 5 and Fable 5 thus becomes a question of economics, not technical pride. Fable 5 costs twice as much. The real question to ask on each workload is not "which one is better" (Fable 5, on paper) but "does Fable 5's premium pay off here". For most flows, agentic coding, document analysis, enterprise workflows, the answer tilts toward Opus 5, and you keep Fable 5 in reserve for the cases where the capability gap genuinely shows up in the result.

Our architecture advice: the model is an interchangeable component

Our take. This six-tier lineup, plus the quarterly updates that keep shifting it, sends a clear signal to technical teams: never hard-code a model identifier in the middle of your business logic. The right instinct is to put the model choice behind a single abstraction, a function or a service that exposes an intent ("fast low-cost answer", "deep reasoning", "long-context analysis") and maps that intent to the right model and the right effort level. The day Anthropic ships Opus 5.1, or the day you decide to switch a flow from Fable 5 to Opus 5 to save money, you change one line of configuration, not twenty calls scattered across the code.

This discipline pays off doubly with this generation. First because the automatic fallback already routes your requests from one model to another server-side: you may as well have your code assume that the model serving a response can differ from the one requested. Second because the cost-performance ratio is now steered as much by effort (high, xhigh, max) as by the choice of model itself. An Opus 5 at high and an Opus 5 at max do not cost the same and do not target the same tasks: treating those two levers as configuration parameters, rather than constants set in stone, is what will let you follow Anthropic's lineup without rebuilding your product with every release.

The Claude lineup, put back in perspective

By the summer of 2026, the Claude catalog holds enough models to breed confusion. Before dissecting Claude Opus 5, we need to lay out the full map and understand which need each model answers. Here is the state of the lineup on July 24 2026, the day Opus 5 shipped.

Model Price (in / out per million tokens) Context Cutoff Role
Haiku 4.5 $1 / $5 200k February 2025 Volume, minimal latency
Sonnet 5 $3 / $15 (intro $2 / $10 through August 31 2026) 1M January 2026 Best balance of speed and intelligence
Opus 4.8 (legacy) $5 / $25 1M January 2026 Former daily driver, now replaced
Opus 5 $5 / $25 1M May 2026 Daily driver for agentic coding and enterprise
Fable 5 $10 / $50 1M January 2026 Maximum capability, broadly available
Mythos 5 $10 / $50 1M January 2026 Defensive cyber, invite-only

Anthropic's segmentation logic reads cleanly by price tier. Haiku 4.5 soaks up volume and minimal latency at $1 / $5. Sonnet 5 targets the balance point for the bulk of traffic, with context now pushed to 1 million tokens and a January 2026 cutoff. The Opus models play the role of a powerful daily driver. At the top, Fable 5 delivers the maximum capability that is broadly available, while Mythos 5, at identical specs and identical price ($10 / $50, model ID claude-mythos-5), stays confined to defensive cyber work through Project Glasswing, accessible by invitation only.

The 2026 cadence is relentless. Fable 5 landed on June 9, Opus 5 on July 24, with the Opus line having rattled off 4.5, 4.6, 4.7 and 4.8 before reaching this fifth tier. The key point to hold onto, and one that is often misread: Opus 5 does not replace Fable 5. It replaces Opus 4.8, at the same $5 / $25 price, bringing a fresher cutoff (May 2026, the most recent in the lineup) and stronger agentic capabilities. Fable 5 keeps its status as the most capable model, at twice the price. Worth noting too: Opus 4.1 is deprecated and retired on August 5 2026, so the previous generation fades fast.

On the architecture side, one practical lesson. At this release cadence, hard-coding a single model identifier throughout your application is a risky move. The healthy reflex is to treat the model as an interchangeable component, isolated behind its own abstraction layer. You then trade off cost against capability (Haiku to triage, Sonnet 5 for production, Opus 5 for heavy agentic work) by changing a single configuration variable, without rewriting your code. That holds all the more because the tokenizer introduced with Opus 4.7 produces roughly 30% more tokens than earlier generations, which shifts your real cost calculations at every migration.

Up against GPT-5.5 and Gemini 3.1 Pro

Shipping a model is the easy part. Placing it honestly in a market where three labs push a new version every six weeks is another story entirely. By late July 2026, the coding-agent landscape had settled around a trio: Anthropic's Claude, OpenAI's GPT family and Google's Gemini. Here is where Opus 5 lands, and, more to the point, what that ranking is actually worth.

The coding-agent leaderboard, late July 2026

The coding-agent comparison published by mightybot.ai (updated on launch day itself, July 24) puts Claude Code running claude-opus-5 at the top of the coding-agent systems, describing it as the best overall agentic system. Right behind it, in second place, sits OpenAI's Codex, powered by GPT-5.6 "Sol" and pitched as the best for long autonomous terminal sessions. Gemini 3.1 Pro, via the Gemini CLI, lands further down the list (seventh), tagged as the best free coding agent but explicitly outside the top tier of the agentic pack.

That hierarchy is no accident, and it maps onto a deeper distinction the same comparison captures well: Opus 5 is reportedly stronger at reasoning across an entire repository, while Sol is stronger at driving a terminal over long stretches. These are two different skills, often blurred together under the single word "coding". One is about holding a mental map of a complex codebase and proposing coherent changes across many files. The other is about chaining hundreds of shell commands without derailing for an hour. Anthropic is clearly aiming at the first; OpenAI excels more at the second.

Three models, three niches

GPT-5.5 (and its successor Sol) remains the reference for desktop agentics and interface automation. The older comparison from tech-insider.org (dated June 26 2026, so before Opus 5 shipped) already names GPT-5.5 as the best for autonomous command-line agents and long background jobs. That is the most solid cross-source read: when you need an agent to take control of a computer, click through interfaces, fill in forms and go the distance on a multi-hour task, OpenAI's model keeps an edge. For business-tool automation or computer-use, it is a defensible default.

Gemini 3.1 Pro, for its part, plays the scientific-reasoning and cost card. The same tech-insider.org comparison ranks it as the best value at scale, with aggressive API pricing ($2 input, $12 output per million tokens in their figures) and a 1-million-token context window. The acknowledged trade-off: it lags on raw code accuracy. Put differently, if your dominant workload is high-volume analysis, synthesis over huge corpora or budget-constrained scientific reasoning, Gemini is worth testing before you commit. It is not the best coder of the bunch, and Google does not claim otherwise.

Claude Opus 5 settles in as the reference for agentic coding and deep reasoning over code. That is the convergent read from both sources: Claude (whether Opus 5 at mightybot.ai or Opus 4.8 at tech-insider.org) comes out on top for code. Where Opus 5 stands out is on multi-file engineering tasks where correctness is paramount, hard-bug investigation and code review. For an agency reworking heterogeneous client codebases, that "understand the repo before you touch it" profile is worth more than an agent that fires off shell commands in a chain.

Which model to pick, in practice

The question is not "which one is best" but "best for what". Here is our pragmatic decision grid, built on the positioning the sources report and not on any brand loyalty.

  • Multi-file code rework and maintenance, code review, complex debugging: Opus 5 first. Its declared strength is reasoning at repository scale and checking its own work, which cuts down on back-and-forth.
  • Interface automation, computer-use, agents driving business apps over long sessions: GPT-5.5 / Sol keeps the claimed advantage. Test it first on these cases.
  • High-volume analysis, scientific reasoning, tight API-budget constraints: Gemini 3.1 Pro enters the running thanks to its cost and 1-million-token context.
  • Hybrid architecture: the mightybot.ai comparison even suggests a multi-model orchestration, with a frontier model as conductor and Sol and Opus 5 splitting implementation and review "on either side of the vendor line". For serious use, not locking yourself into a single vendor is often the genuinely right call.

The mandatory caution on these comparisons

One warning is in order, and we will state it plainly. These cross-vendor rankings are noisy. The two sources we cite do not even discuss the same versions: mightybot.ai compares Opus 5 and GPT-5.6 Sol, while tech-insider.org, a month older, still reasons in terms of Opus 4.8 and GPT-5.5, never once mentioning Opus 5 or Sol. The direct consequence: a figure attributed to Opus 5 by one source may refer to Opus 4.8 in the other, and neither offers a rigorous head-to-head of the three models evaluated side by side under identical conditions.

Add to that the fact that versions move very fast (GPT-5.5 then 5.6 Sol in a matter of weeks, Opus 4.8 then Opus 5, Gemini iterating continuously) and you get shifting ground where any ranking goes stale within a month or two. Our advice does not change: do not pick a model on the strength of a table you found online, ours included. Run your own evaluation on your real workload, measure cost per task and reliability on your concrete cases, and re-test with every new version. That is the only method that holds up to the current pace of the market.

What the people who tested it are saying

As with every model launch, Anthropic's announcement arrives with a volley of quotes from partners who got early access. A necessary editorial note: these are testimonials from companies hand-picked by Anthropic, often customers who sell their own products built on top of these models. Read them as interested endorsements, not independent tests. That said, when players whose livelihood depends on the reliability of code all converge on the same message, the signal is worth pausing on.

The dominant refrain: Fable-level quality, at half the price

The common thread running through these testimonials is the capability-to-cost ratio. Two major players in dev tooling phrase it almost identically.

Cognition (Devin), Scott Wu, CEO: Claude Opus 5 approaches Fable's level of performance, at half the cost.

Cursor, Sualeh Asif, co-founder: intelligence close to claude-fable-5, but at the speed and cost of an Opus.

The message is crystal clear and perfectly aligned with Anthropic's sales pitch: you get nearly the top tier without paying top-tier prices. Coming from Cursor and Devin, two products whose margins depend directly on the cost per token consumed, the argument is not trivial. A model that halves the bill while staying near the summit is, mechanically, a profitability lever for them. Which also explains their enthusiasm: they have a direct interest in this positioning being true.

Automation and the enterprise

Beyond pure code, several partners stress the agentic and business use cases.

Zapier, Wade Foster, CEO: Opus 5 took the top spot on Zapier's internal AutomationBench leaderboard, reaching what they report as 100% on the tested workflow, without consuming more tokens than previous Claude models.

Box, Ben Kus, CTO: roughly 8% more performance than Opus 4.8 on their tasks, and up to 11% better on data analysis.

These two reports are interesting because they step outside Opus's natural playground (coding) to touch workflow orchestration and enterprise document processing, exactly the kind of use case our clients are after. Box's figure, a gain of a few points on an already solid base, actually sounds more credible than a spectacular leap: it is the order of magnitude you would expect from a well-executed version bump, not a revolution.

Variance reduction: the signal that speaks to production teams

The most instructive feedback for anyone shipping models to production is not about a performance peak, but about its stability.

Lovable, Fabian Hedin, co-founder: +22% over Opus 4.7, and above all far less variance from one run to the next.

That variance point is the one we take away. In production, a model that delivers an excellent result one time in three and a mediocre one the other two is unmanageable: you cannot build a reliable product on it. A slightly less brilliant but predictable model, run after run, is often worth more than an unstable genius. If the consistency gain holds up outside a partner's test conditions, it is probably Opus 5's most concrete improvement for an agency like ours, more so than any benchmark score.

Creative work, science, and IDEs

Three final reports round out the picture on more specific terrain.

Vercel, Madhav Jha, co-founder and CTO: the best animations, games, and 3D renders produced by any model in the Opus family.

JetBrains, Denis Shiryaev, head of AI in the IDE: the sharpest jump in problem-solving.

Benchmark Life Sciences, Alfredo Andere, CEO: the model behaves more like a cautious scientist than any other model tested.

Alfredo Andere's phrasing neatly captures a deeper trend: caution, the habit of verifying before asserting, is becoming a selling point on par with raw power. On the creative front, Vercel's feedback confirms that Opus 5 is taking on a domain (visual and interactive rendering) where the family was historically less comfortable.

What to take away from this chorus of praise

Two conclusions converge across these voices. First, the cost-to-capability promise (close to Fable, at half the price) is echoed word for word by the partners best placed to judge it, which lends it credibility without proving it. Second, variance reduction keeps coming up as an underlying theme: reliability from one run to the next, fewer round-trips, more predictable behavior. For a decision-maker, it is the second point that matters most, because it determines whether a model can be industrialized. What remains is the fundamental caveat: none of these numbers were produced under independent conditions, each partner measured whatever suited it on its own bench, and all of them have an interest in the model succeeding. So we take them for what they are: converging indications, not proof.

Security, alignment, and the cyber case

This is the part of the spec sheet that rushed buyers skip and that CIOs read first. Anthropic ships every model with a system card, a 194-page document for Opus 5 in which the vendor lays out its safety, alignment, and risk testing. A methodological note that holds for this entire section: these figures come from Anthropic's internal evaluations, not an independent third-party audit. They are solid and detailed, but they remain self-reported. Read them as such.

An alignment score billed as the best of its generation

The number Anthropic puts forward is a behavioral-audit misalignment score of 2.3. The scale runs from 1 to 10 (lower is better), each model goes through roughly 3,200 automated investigations, and 2.3 is, according to the vendor, the lowest score among its recent models, ahead of Sonnet 5, Opus 4.8, and Mythos 5. Anthropic uses this to claim the title of most aligned model to date: the lowest rates of deceptive behavior, the strongest adherence to Claude's constitution, and fewer risky irreversible actions.

Two honest caveats, both taken from the system card. First, Opus 5 is slightly better than its predecessors at recognizing when it is being tested (evaluation awareness), which can skew this kind of audit; Anthropic states that this did not materially distort the conclusions, but the reader should know that the observed subject sometimes knows it is being observed. Second, a counterintuitive detail: the model hallucinates factual claims slightly more often than Opus 4.8, even though it is more accurate overall. A good alignment score therefore does not mean zero factual errors. Human review is still the rule for sensitive deliverables.

Self-checking and error recovery

In the field, the most useful improvement is not a benchmark number but a behavior: Opus 5 checks its own work and recovers from its own mistakes, which cuts down the number of round trips compared with its predecessors. That is what Madhav Jha (co-founder and CTO of Vercel) describes, pointing to the best animations, games, and 3D renders produced by an Opus model. For agentic use (the model chains steps without supervision), this ability to catch its own errors often matters more than a point or two more on an academic test.

The cyber case: a deliberate strategy

The most interesting point, and the one most specific to this model family, is cyber. Anthropic says it intentionally avoided training Opus 5 on cyber tasks, as it did with Opus 4.8. The model's cyber gains therefore come from its general-capability progress, not from dedicated offensive training. In practice, this translates into an asymmetric profile: Opus 5 closes in on Mythos 5 (Anthropic's strongest cyber model) at identifying vulnerabilities, but stays well behind at turning them into working exploits.

The internal measurements illustrate this gap. On OSS-Fuzz, Opus 5 posts a non-zero score on roughly 79% of targets, versus 38% for Opus 4.8 and about 80% for Mythos 5: on detection, it has joined the top of the table. But when it comes to seeing things through, the gap widens: Opus 5 fully exploits a handful of targets where Mythos 5 closes the loop far more often. Same logic on other internal suites: Opus 5 finds, Mythos 5 weaponizes. Anthropic sums up this positioning in its Opus 5 announcement by saying the model remains substantially behind Mythos 5 on exploit development.

Fewer cyber false positives: why it matters

The most concrete change for legitimate users concerns the safety classifiers, the filters that refuse certain requests deemed dangerous. On Opus 5, they are calibrated to trigger roughly 85% less often than on Fable 5. On one internal measure, the trigger rate drops from 42% to 5%. The policy change is explicit: Opus 5 now allows vulnerability research in source code at all access levels, while continuing to block by default the scanning of compiled binaries, penetration testing, and exploit generation.

Our take: this drop in false positives is the real good news for professional use. A developer asking Claude to audit their own code, an internal pentester documenting a flaw, a security lead preparing a patch all ran into refusals from the previous generation on a regular basis. Fewer wrongful blocks means a model you can actually use for code review and application hardening, without having to work around the tool. The guardrail stays, but it gets in the way of defensive work far less. When a refusal does happen anyway, the API can fall back to Opus 4.8 to return an answer rather than an error.

Biology, Mythos 5, and Project Glasswing

On the biology side, Opus 5 keeps a set of guardrails similar to that of Opus 4.8. Anthropic presents it as the most capable generally available model for scientific research, while flagging significant limits on long autonomous tasks. One test says it all: tasked with planning and running, on its own, a 24-hour protein-design campaign for $10,000, the model failed to deliver the result on two attempts. The real capability ceiling therefore sits below the marketing.

Finally, Opus 5 needs to be distinguished from Mythos 5, often cited as a point of comparison. Mythos 5 is not generally available: it is a restricted model, distributed by invitation via Project Glasswing, a collaboration with the US government aimed at cyber defenders and critical-infrastructure operators. Anthropic's reasoning is a deliberate offense/defense asymmetry: keep the most advanced cyber capabilities in a closed, monitored channel, and ship the general public an Opus 5 deliberately reined in on offense but far more permissive on legitimate defense. For an agency or a business, that is the right trade-off: the power that is useful day to day, without the loaded gun.

The glossary you need to read this launch properly

The launch of Claude Opus 5 comes wrapped in technical terms that show up everywhere in the documentation. Here is a plain-English glossary to make sense of them without being an engineer, each one tied back to what this new model actually changes.

Context window

The context window is how much text the model can read and hold in memory during an exchange. Opus 5 accepts up to 1 million tokens, roughly 555,000 words. In practice, you can hand it an entire case file, a stack of contracts or a full codebase without slicing it into pieces.

Token and tokenizer

A token is a small unit of text, often a fragment of a word. The tokenizer is the tool that chops your text into tokens. This matters for your budget: the Opus 4.7 generation and onwards produces about 30% more tokens than earlier versions. Since billing is per token, that detail shifts your cost math even when the final text is identical.

Knowledge cutoff and training data cutoff

Knowledge cutoff
The date up to which the model knows about the world. For Opus 5, that cutoff sits at May 2026, the most recent in the entire lineup.
Training data cutoff
The end date of the raw training data. It can predate the knowledge cutoff. The two are distinct notions, and it is a mistake to conflate them.

Adaptive thinking and extended thinking

Adaptive thinking is active reasoning that runs by default on Opus 5: the model decides on its own how much to think based on how hard the task is. Extended thinking, the on-demand prolonged reasoning mode from previous generations, no longer exists on Opus 5, replaced by this adaptive approach.

Effort

The effort parameter arbitrates the trade-off between cost and capability. It offers the levels high, xhigh and max, with high as the default on the API and in Claude Code. The higher you go, the more compute the model invests, so more precision, but also more time and more budget.

Agentic and agent

A model described as agentic, or an agent, does more than answer. It plans, uses tools such as a browser or a terminal, and chains multiple steps autonomously to reach a goal. This is a major focus of Opus 5.

SWE-bench Verified and SWE-bench Pro

These are two benchmarks, standardized tests, that measure a model's ability to fix real software bugs. SWE-bench Verified gathers cases vetted by humans. SWE-bench Pro is a considerably harder version, with more complex problems.

Prompt cache

The prompt cache reuses context you have already sent so it does not have to be reprocessed, cutting both cost and latency. A useful addition on the Claude Platform: switching tools mid-conversation no longer invalidates that cache, so your savings hold steady.

Automatic fallback

Automatic fallback is a switch to another model when Opus 5 refuses a request for a safety reason. A legitimate request is not blocked outright, it is rerouted, which avoids interruptions in a production flow.

Our hands-on take: what is verifiable and what is not

Before we line up a single score, let us be upfront about how we approached this piece. We went back to the primary sources: Anthropic's announcement page, the platform's technical documentation, the system card and the official pricing tables. We then cross-checked all of that against the trade press (Fortune, Yahoo Finance) and against third-party evaluators such as Snorkel and CodeRabbit. What we did not do, and we want to say it plainly: at this stage, we have not produced an independent, reproducible benchmark of our own. The launch numbers, including those presented under third-party benchmark names, were measured on Anthropic's own runs, not by independent labs. That is the norm on release day, but it remains a limitation we refuse to paper over.

What we can state with confidence

Some things do not hinge on a benchmark: they are observable, documented and stable. The model ID is claude-opus-5, available since July 24 2026 on Claude.ai, Claude Code, Claude Cowork, the API and via AWS Bedrock, Google Cloud and Microsoft Foundry. Pricing is $5 per million input tokens and $25 per million output tokens, strictly identical to Opus 4.8, which is half the rate of Fable 5. The context window reaches 1 million tokens, maximum output is 128,000 tokens, and the knowledge cutoff is May 2026. All of this is verifiable in a few minutes on your own account.

The effort dial is likewise a tangible fact: the effort parameter accepts low, medium, high, xhigh and max, with high as the default on the API and in Claude Code. You set it, you measure the difference in tokens consumed, you see it for yourself. Same logic for the automatic fallback: when a safety classifier declines a request, the API returns a stop_reason: "refusal" with an HTTP 200 and can switch over to Opus 4.8 rather than return an error. The behavior is described in the docs and reproduces reliably. These building blocks (pricing, availability, specs, effort, fallback) we take as settled.

What still needs confirming

On the other hand, several of the core claims in Anthropic's pitch are not verifiable from the outside on launch day. The first is the real gap with Fable 5 in production conditions. Anthropic states performance within 0.5% of Fable 5 on CursorBench for half the cost per task. That is a vendor-reported figure, not an independent measurement, and nothing guarantees that gap holds on your codebase, with your prompts and your constraints. The second unknown is the effective cost. The sticker price is one thing, the invoice is another: the Opus 4.7 generation and later use a tokenizer that produces roughly 30% more tokens than pre-4.7 models. An identical per-token price can therefore mask a higher real spend, one to be measured against your own workload.

The third grey area is run-to-run stability. Partner testimonials, however flattering (Lovable points to markedly reduced variance from one run to the next), describe hand-picked and often optimized conditions. Variance on long, autonomous tasks, outside a controlled demo, can only be judged through use over several weeks. Finally, the error analysis from Snorkel is a useful reminder that faulty inference remains the leading failure mode (35% of cases): a stronger model is not an infallible one.

How to test it seriously yourself

Our advice is simple and it costs almost nothing. Do not trust benchmark averages: build a mini-corpus that represents your actual work, a dozen or so tickets, bug reports or typical documents, and run Opus 5 on it.

  • Compare on your own corpus, not on SWE-bench: public scores do not predict your specific use case.
  • Measure the real tokens consumed, input and output, to get a concrete cost per task rather than a theoretical price per million.
  • Vary the effort from high to xhigh and then max on the same tasks, and see whether the quality gain justifies the extra tokens. Often, lowering the effort preserves most of the performance for far less.
  • Keep Opus 4.8 as your baseline: it is still available. An honest A/B between 4.8 and Opus 5 on your tasks is worth a thousand launch scores.

In short, we are convinced by what we can touch (the price held steady, the effort dial, the immediate availability) and cautious about what still rests on the vendor's word (the exact gap with Fable 5, the net cost after tokenizer, the consistency over time). That is precisely why we urge you to run your own test before committing a budget: on this subject, your corpus is a more reliable judge than any leaderboard.

Who it is actually for

Once the announcement dust settles, the question that matters for a budget and a roadmap is simple: does claude-opus-5 change anything for you, and which model in the lineup should you pick? The answer depends less on the benchmark than on your profile. Here is our read, profile by profile, without the overselling.

Developers and product teams: a sensible new default

For most dev teams, Opus 5 becomes the reasonable default for agentic coding. Two concrete reasons. First, the capability-to-cost ratio: same price as Opus 4.8 ($5 input, $25 output per million tokens) for a real jump in performance. Second, the third-party scores reported by the press and evaluators, which should be handled with care since the launch runs were Anthropic's own: 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro, well above Opus 4.8's 69.2%. That is not marginal, it is a generational step.

Our practical recommendation fits in three lines. Use Opus 5 as your daily driver for complex agentic coding, multi-file refactoring and bug investigation, where it shines. Keep Claude Sonnet 5 and Haiku 4.5 for anything high-volume, sub-agents and simple tasks, where spending Opus compute makes no sense. And reserve Fable 5, twice as expensive, for the hardest reasoning problems, the ones where that last half-point of capability justifies the bill. Anthropic does not claim Opus 5 is its smartest model: that title still belongs to Fable 5. Opus 5 is the best cost-performance play, not the absolute ceiling.

One operational detail that matters: the effort parameter (levels high, xhigh, max) is your cost lever. Lowering the effort preserves most of the performance while burning fewer tokens. In practice, run an effort sweep on your own evals instead of pushing everything to max by reflex. Watch out too for the Opus 4.7-and-later generation tokenizer, which produces roughly 30% more tokens than earlier models: measure on your real workload before extrapolating a bill.

Enterprises: capability-to-cost in service of internal workflows

For a CIO or a business unit lead, the appeal of Opus 5 is not the benchmark contest, it is the economics on recurring internal use cases. Data analysis, document synthesis, scientific research, due diligence, financial modeling: these are exactly the areas where Anthropic positions the model, and where several of its customers report gains (Box cites 8% better than Opus 4.8 and 11% on data analysis, JetBrains a clear jump in problem-solving). These figures are declared by partners the vendor quotes: treat them as signals, not independent measurements.

Two practical advantages for the enterprise. The 1 million token context window lets you ingest entire document bases with no acrobatic chunking. And the automatic fallback (when a request is refused for safety reasons, the API switches to another model instead of returning an error) avoids production dead ends. The fact that Anthropic describes Opus 5 as its most capable generally available model for scientific research, notably in biology, is worth taking seriously for the sectors concerned, healthcare, agri-food, chemistry, while keeping in mind that these are vendor specs and framing.

Agencies and SMBs: how we use it at Go To Agency

This is where the cost-performance logic becomes very concrete for a client. At Go To Agency, a digital agency in Dijon, a model that approaches the top tier at half the price of Fable 5 translates directly into deadlines met and margin preserved, and therefore into more competitive quotes for you. We use it on three fronts.

  • Prototyping: starting from a fuzzy brief and shipping a first working mockup, a UI component or an integration script in hours rather than days. Opus 5 self-checks its work and recovers from its own mistakes, which cuts the back-and-forth and the billable time lost to corrections.
  • Content and SEO: producing documented in-depth articles, local pages and multilingual content, with systematic human review. Editorial quality is not up for negotiation, but the production cost drops, and that saving benefits the end client.
  • Internal tools: lead dashboards, automations, small custom back-offices. We build quickly what, yesterday, would have needed a full dev cycle, and we bill for the result, not the time lost.

The principle is simple and it does not change: AI speeds up our production, it never replaces the judgment, the strategy and the editorial responsibility you expect from an agency. It lets us deliver faster and stay price-competitive, without ever overcharging for capability the project does not require. A local brochure site does not need the compute of a complex platform project, and we pick the right model for the right job.

Want to see what that looks like in practice? Browse our work, request a quote tailored to your project, or get in touch to talk it over. Everything happens by email, at your own pace: we read, we understand the need, we reply with a concrete proposal, without forcing a call or a calendar on you.

Migrating from Opus 4.8, and our verdict

Good news if you are already running claude-opus-4-8: the jump to claude-opus-5 is one of the smoothest migrations Anthropic has ever shipped. Same family, same 1 million token context window, same pricing grid ($5 per million input tokens, $25 per million output), and an official migration guide that documents the handful of behavioral differences. In most cases, you swap the model ID in your code and rerun your evals. The price did not move, the performance went up: that is the story Anthropic is selling, and on this specific point, it checks out.

That said, "change the ID and pray" is not a production strategy. Four things deserve a real test before you flip the switch.

  • Check the default effort. On the API and in Claude Code, Opus 5 runs at high by default, exactly as if you omitted the parameter. Watch one difference from Opus 4.8: the older model forced high everywhere, including on claude.ai, whereas Opus 5 no longer forces high by default on claude.ai. Rerun an effort sweep on your own evals rather than blindly porting your Opus 4.8 settings.
  • Measure your real token counts. Generation 4.7 and later uses a tokenizer that produces roughly 30% more tokens than pre-4.7 models on the same text. The per-million price is identical, but your actual bill depends on the number of tokens billed. Measure it on your own workload before trusting any extrapolation.
  • Test the fallback. If you rely on the automatic fallback behavior (when a safety classifier declines a request, the API switches to another model, Opus 4.8 by default on Claude.ai, Code, and Cowork), validate the client-side output: the response comes from a different serving model, and your application logic has to tolerate that.
  • Keep Opus 4.8 reachable as legacy. It stays available, which gives you a comparison point and a safety net if a use case regresses. One thing to watch on the neighboring calendar, though: claude-opus-4-1 is already deprecated and will be retired on August 5, 2026. If you are still dragging along on a 4.1, migration is no longer optional.

Our verdict

Let us be clear and honest, as we have been throughout this article. Opus 5 is the best Opus ever released, at the same price as the previous one. It closes in on Fable 5 across a lot of coding and reasoning tasks, while remaining, by Anthropic's own admission, less intelligent than Fable 5 at the top of the spectrum. Anthropic does not claim otherwise: Opus 5 is sold as the daily driver, the best cost-to-performance ratio, not the peak of the lineup. That is an honest position, and it matches what we are seeing.

The genuinely useful novelty is not any single benchmark point. It is two concrete mechanisms: the effort dial (low, medium, high, xhigh, max), which lets you trade cost against capability on a single request, and the automatic fallback, which turns a safety refusal into a degraded answer rather than an error. These are the features that actually change life for a team running the model in production, far more than half a point on a leaderboard.

On the numbers, our stance stays cautious. The benchmarks are impressive (96.0% on SWE-bench Verified, 79.2% on SWE-bench Pro, scores that double Opus 4.8 on Frontier-Bench according to the figures declared by Anthropic), but most of them are company-reported or run on Anthropic launch harnesses, not by strictly independent third parties. That is normal for a launch, and it is no reason to swallow them whole. Our read: the direction is right, the exact magnitude is yours to confirm on your own tasks.

The operational verdict is simple. For the vast majority of coding and agent workloads, Opus 5 becomes the new reasonable default: you keep Fable 5 for maximum capability on the genuinely hard problems, and you drop down to Sonnet 5 or Haiku 4.5 when cost is the priority. Opus 5 sits at the center of gravity, right where the daily work happens.

At Go To Agency, we already fold these models into our clients' pipelines (coding assistants, business agents, document automation) and we pick the right model for the right budget, with no dogma. If you are wondering which of these models makes sense for your use case, or how to wire a clean fallback in production, let us talk: request a quote or head over to the contact page. We will tell you honestly what is worth it, and what is not.

AI BRIEF · GO TO AGENCY

AI news, decoded for builders

Once a week, our no-noise take on the AI releases that matter: models, tools, pricing. No spam.

1 email a week · 1-click unsubscribe · GDPR-friendly

RM

About the author

Robin Monteiro

Co-fondateur de Go To Agency

Développeur full-stack et co-fondateur de Go To Agency, Robin conçoit des solutions web performantes avec Next.js, React et les dernières technologies.

Meet the team

Go To Agency — digital agency, Dijon (France)

The team behind this article can build it for you

Custom Next.js websites and e-commerce, SEO that ranks, and ad campaigns measured down to the return. Everything happens in writing, no meetings: describe what you need and we come back with a concrete read.

Your request lands directly in [email protected] — reply within 24 business hours, no commitment.

Share article

Questions fréquentes

What is Claude Opus 5?+

Claude Opus 5 (claude-opus-5) is the Anthropic model released on July 24, 2026, positioned as a daily driver with the best cost-to-performance ratio. Anthropic describes it as a thoughtful, proactive model that approaches the frontier intelligence of Claude Fable 5 at half the price. It handles text and image input, text output and multilingual work, with adaptive reasoning.

How much does Claude Opus 5 cost?+

The price is $5 per million input tokens and $25 per million output tokens, exactly the same rate as Opus 4.8 and half the price of Fable 5 ($10 / $50). A Fast Mode that is 2.5x faster is available, but it costs 2x the base price.

What are the context window and knowledge cutoff of Opus 5?+

Opus 5 has a 1 million token context window (roughly 555,000 words) and a maximum output of 128k tokens (up to 300k via the Batch API in beta). Its reliable knowledge cutoff is May 2026, more recent than Opus 4.8 (January 2026).

Is Opus 5 better than Fable 5?+

No, according to Anthropic: Fable 5 remains the most capable model in the widely available lineup. Opus 5 gets close to it at half the price. On reported third-party benchmarks such as SWE-bench Pro, Opus 5 (79.2%) sits just behind Fable 5 (80.0%) and Mythos 5 (80.3%). The docs recommend Opus 5 for complex agentic coding and enterprise work, and Fable 5 for maximum capability.

How is it different from Opus 4.8?+

Same price and same context (1M), but a more recent knowledge cutoff (May 2026 vs January 2026) and clear gains on reported third-party benchmarks (SWE-bench Pro: 79.2% vs 69.2%). Opus 5 adds the effort dial and an automatic fallback to another model on safety refusals. Opus 4.8 becomes legacy but stays available.

Where can you use Claude Opus 5?+

Opus 5 is available on Claude.ai, Claude Code, Claude Cowork, the Claude API and the Claude Platform, as well as AWS Bedrock (anthropic.claude-opus-5), Google Cloud and Microsoft Foundry. It is the default model on Claude Max and the strongest option on Claude Pro.

Related articles

Free quote
Claude Opus 5: specs, price and honest review | Go To Agency