Comparison7 min readPublished on 2026-07-24

Opus 5 vs Opus 4.8: what really changes and whether it's worth switching

A comparison of Claude Opus 5 and Opus 4.8: same list price, better performance and the new effort dial. When it's worth switching and what to review in your setup.

In a nutshell

Opus 5 costs exactly like Opus 4.8 ($5/$25 per million tokens) but delivers more: SWE-bench Pro from 69.2% to 79.2%, plus the new effort dial that lowers the effective cost per task. For anyone already using Opus 4.8 the switch is almost always worth it. The only things to review: model routing and agents in production.

The difference in one line

Same price, more capability, more control over cost. Opus 5 costs like Opus 4.8 — $5 per million input tokens, $25 on output — but delivers more on almost every front and adds a lever that wasn't there before: the effort dial.

If you've got two minutes and already use Opus 4.8, the short answer is: switch. Below you'll find the numbers and the few things to check first. If instead you're starting from scratch, you're better off reading what Opus 5 is first.

The numbers head to head

On agentic coding, the ground that weighs most today, the jump is clear. On SWE-bench Pro Opus 5 scores 79.2% against Opus 4.8's 69.2%: a full ten points. For scale, Sonnet 5 on the same test sits at 63.2%. On ARC-AGI 3, the benchmark on never-seen-before problems, Opus 5 pulls away from the field by about three times.

On the safety front too there's an improvement: Anthropic calls Opus 5 its "most aligned" Opus, with the lowest score on the internal behavioral audit. For anyone putting agents into production that's no detail: it's less risk surface. (The benchmark figures beyond SWE-bench Verified come from third-party reports in the first days after launch; the order of magnitude is solid, the individual numbers should be taken as indicative.)

The effort dial changes the economics

The list-price comparison tells half the story. The other half is how much you actually spend to get a task done. Here Opus 5 introduces the effort dial: a knob that adjusts how hard the model reasons.

With Opus 4.8 you paid full power on every call. With Opus 5 you raise the effort where the task is hard and lower it where a quick answer is enough, keeping most of the quality and consuming fewer tokens. Result: at the same list price, the average cost per task drops. That's the main reason the comparison isn't "same as before but a bit better," it's a change in the equation.

Want to review your model mix after Opus 5?

30 minutes to discuss your specific case.

Book a call

When it's worth switching (and when to wait)

It's worth switching right away if: you already use Opus 4.8 in production, you run agents, or you have flows where cost per task matters. Same price, more output, more control. There's almost no reason to stay behind.

It makes sense to move more slowly only if you have pipelines heavily tuned and validated on Opus 4.8, with automated evals you can't re-run quickly. In that case it's not "don't switch," it's "switch after re-running your evals." The new model changes the answers: what was tuned before needs re-verifying, not taking on faith.

What to review in your setup

Two things, concrete. First, model routing. With the effort dial you can move onto Opus 5 (at low effort) tasks you used to hand to a smaller model to save money. Redo the math: you might simplify the architecture and raise quality at the same cost.

Second, agents in production. Opus 5 at high effort closes tasks Opus 4.8 was failing; at low effort it saves you money on the simple ones. Re-measure success rate and cost per task on your real flows, not on public benchmarks. And if you use the API, look at the two new betas — mid-conversation tool changes and automatic fallbacks — which cut costs and interruptions. We cover them in Opus 5 for coding.

The full model picture

Opus 5 doesn't live alone: it fits with Sonnet 5 for volume and with the frontier models for extreme tasks. If you're redesigning your model strategy, the point isn't "which is best" but "which task to which model, at which effort." For an up-to-date overview see Sonnet, Opus and Haiku compared and the deep-dive on Claude models.

At Maverick AI we help companies make exactly this choice, without getting attached to a single model. If you want to review your mix after Opus 5, get in touch.

FT
Federico Thiella·Founder, Maverick AI

Works with European companies on Claude and Anthropic ecosystem adoption. Has led AI implementations in private equity, consulting, manufacturing and professional services.

LinkedIn

Want to review your model mix after Opus 5?

We help you re-map which task goes to which model and at which effort, with measurement on real cost and quality. Let's talk.

Write to us

Frequently asked questions: Opus 5 vs Opus 4.8

No, the list price is identical: $5 per million input tokens and $25 on output. The effective cost per task with Opus 5 actually tends to drop, thanks to the effort dial that has you consume fewer tokens on simple tasks.
On agentic coding the jump is about ten points: SWE-bench Pro goes from Opus 4.8's 69.2% to Opus 5's 79.2%. On ARC-AGI 3 the margin is even wider. It's also the most aligned Opus model from a safety standpoint.
Two things: model routing (with the effort dial you can consolidate more tasks onto Opus 5) and validation of agents in production. A new model changes the answers, so it pays to re-run your evals instead of taking the previous setup on faith.
At launch Opus 4.8 stays available via API at the same price. But since Opus 5 costs the same and delivers more, for most uses there's no reason to stay on 4.8. Always check the model lifecycle in the Anthropic documentation before critical pipelines.
Yes, and it's automatic: Opus 5 is already the default on Claude Max and the strongest model on Pro. You don't need to do anything but start using it. The reasoning about routing and effort mainly concerns API users.

Stay informed on AI for business

Get updates on Claude AI, business use cases and implementation strategies. No spam, just useful content.

Want to learn more?

Contact us to find out how we can help your company with tailored AI solutions.

Anthropic implementation partner in Italy. We work with companies in PE, pharma, fashion, manufacturing and consulting.

Related articles

News

Claude Opus 5: the flagship model now has a "dial" for cost — what changes for businesses

Anthropic has released Claude Opus 5: close to Fable 5 on many tasks at half the cost per task, with an "effort dial" that adjusts how hard the model reasons. Pricing, benchmarks and what it means for using AI in your company.

Guide

Claude Opus 5 for coding: benchmarks, effort dial and what's new for agent builders

Claude Opus 5 for development: SWE-bench Pro at 79.2%, CursorBench close to Fable 5 at half the cost, the effort dial and the two API betas for agents. What changes in Claude Code.

News

Claude Sonnet 5: Opus-Level Performance at Sonnet Pricing — What Changes for Businesses

Anthropic has released Claude Sonnet 5: more agentic, close to Opus 4.8, and cheaper. Pricing, performance, and what it means for companies using AI.

Technical

Claude Sonnet vs Opus vs Haiku: Which Model to Use

Detailed comparison of Claude's three model tiers: Haiku, Sonnet and Opus. Strengths, use cases, pricing and decision framework to help you choose the right Claude model for each task.

News

Claude Opus 4.7: All the new features of April 16, 2026

Claude Opus 4.7 released April 16, 2026: hybrid reasoning model with record benchmarks in coding, vision and legal. Pricing, availability and what changes for enterprises.

Strategy

How much does Claude AI cost for businesses: pricing, plans and real costs

Complete guide to Claude AI costs for businesses. API pricing by model, enterprise plans, typical integration project costs and how to optimize spending.

Book an introductory call
Opus 5 vs Opus 4.8: comparison, pricing, upgrade (2026) | Maverick AI