The difference in one line
Same price, more capability, more control over cost. Opus 5 costs like Opus 4.8 — $5 per million input tokens, $25 on output — but delivers more on almost every front and adds a lever that wasn't there before: the effort dial.
If you've got two minutes and already use Opus 4.8, the short answer is: switch. Below you'll find the numbers and the few things to check first. If instead you're starting from scratch, you're better off reading what Opus 5 is first.
The numbers head to head
On agentic coding, the ground that weighs most today, the jump is clear. On SWE-bench Pro Opus 5 scores 79.2% against Opus 4.8's 69.2%: a full ten points. For scale, Sonnet 5 on the same test sits at 63.2%. On ARC-AGI 3, the benchmark on never-seen-before problems, Opus 5 pulls away from the field by about three times.
On the safety front too there's an improvement: Anthropic calls Opus 5 its "most aligned" Opus, with the lowest score on the internal behavioral audit. For anyone putting agents into production that's no detail: it's less risk surface. (The benchmark figures beyond SWE-bench Verified come from third-party reports in the first days after launch; the order of magnitude is solid, the individual numbers should be taken as indicative.)
The effort dial changes the economics
The list-price comparison tells half the story. The other half is how much you actually spend to get a task done. Here Opus 5 introduces the effort dial: a knob that adjusts how hard the model reasons.
With Opus 4.8 you paid full power on every call. With Opus 5 you raise the effort where the task is hard and lower it where a quick answer is enough, keeping most of the quality and consuming fewer tokens. Result: at the same list price, the average cost per task drops. That's the main reason the comparison isn't "same as before but a bit better," it's a change in the equation.
Want to review your model mix after Opus 5?
30 minutes to discuss your specific case.
When it's worth switching (and when to wait)
It's worth switching right away if: you already use Opus 4.8 in production, you run agents, or you have flows where cost per task matters. Same price, more output, more control. There's almost no reason to stay behind.
It makes sense to move more slowly only if you have pipelines heavily tuned and validated on Opus 4.8, with automated evals you can't re-run quickly. In that case it's not "don't switch," it's "switch after re-running your evals." The new model changes the answers: what was tuned before needs re-verifying, not taking on faith.
What to review in your setup
Two things, concrete. First, model routing. With the effort dial you can move onto Opus 5 (at low effort) tasks you used to hand to a smaller model to save money. Redo the math: you might simplify the architecture and raise quality at the same cost.
Second, agents in production. Opus 5 at high effort closes tasks Opus 4.8 was failing; at low effort it saves you money on the simple ones. Re-measure success rate and cost per task on your real flows, not on public benchmarks. And if you use the API, look at the two new betas — mid-conversation tool changes and automatic fallbacks — which cut costs and interruptions. We cover them in Opus 5 for coding.
The full model picture
Opus 5 doesn't live alone: it fits with Sonnet 5 for volume and with the frontier models for extreme tasks. If you're redesigning your model strategy, the point isn't "which is best" but "which task to which model, at which effort." For an up-to-date overview see Sonnet, Opus and Haiku compared and the deep-dive on Claude models.
At Maverick AI we help companies make exactly this choice, without getting attached to a single model. If you want to review your mix after Opus 5, get in touch.