Effort is the dial that matters in Opus 5
AI / GenAI·6 min·25 July 2026·Claude Opus 5 — Part 2 of 4

Effort is the dial that matters in Opus 5

Claude Opus 5 costs exactly what its predecessor cost. Five dollars per million input tokens, twenty-five per million out. That sounds like a non-story until you work out what it means: if the price per token is fixed, your bill is decided by the number of tokens. Which is precisely what the effort dial controls.

In this release that is no longer a detail. It is the main story.

What effort is

Effort is a setting with five positions: low, medium, high, xhigh and max. It governs how deeply the model reasons and how many tokens it may spend on a job before answering. The default is high.

You meet it in different places. On the API you set it as a parameter. In Claude Code, xhigh is the default. If you use Claude through the app, the dial is not on your screen, and the model decides per turn how much thinking a question deserves.

What is interesting is not that the dial exists. Opus 4.8 had it too. What is interesting is that the settings mean something different this time.

The surprise sits at the bottom end

The habit until now was: set it high, or the model gets dumb. On Opus 4.8 that was roughly true.

On Opus 5 it is not. The clearest number comes from Zapier’s AutomationBench, which measures whether a model can carry business tasks from start to finish. On its lowest setting, Opus 5 passes more tasks there than any other model does on its highest. Not more than Opus 4.8. More than every other model.

Anthropic says it outright in their own guidance: use low and medium liberally where your evals show quality holds. And: if you carried your effort defaults over from a previous model, run the sweep again. That is unusually direct advice from a vendor that earns per token.

Which setting for what

This is how I divide it up now. Not a law, but a starting point.

SettingForWhat you notice
lowSummarising, classifying, short questions, anything speed-sensitiveFast and cheap, at quality that used to need high
mediumMost daily work: writing, analysis, one-off code questionsThe new default if you watch cost
highThe default value. Precision work with a contained scopeWhat you were used to, for fewer tokens
xhighCoding and agentic work. The recommended starting point thereMultiple files, longer chains, less giving up
maxThe genuinely hard cases where correctness beats costBest result, clearly pricier, occasionally overthinks

What stands out in that table: there is no longer any reason to sit on high by default without thinking about it. Down is a win, up is a win, standing still is the only option that gains you nothing.

What the numbers say about the bill

Almost every measurement customers have published points the same way: not better answers for more money, but the same or better answers for less.

A financial customer measures nine percentage points higher accuracy across effort levels, with a third fewer turns and tool calls and sixty percent less time. A legal customer reaches comparable performance using twenty-six percent fewer tokens than Opus 4.8 at max reasoning. On CursorBench the model at max lands within half a percent of Fable 5, at half the cost per task.

Translate that into your own situation and you get a different conversation than at previous launches. Not “is the upgrade worth it”, but “where do I have my settings too high”.

The caveat nobody prints alongside

Two things bother me.

Effort makes your costs less predictable, not more. You now have a dial you can steer with, but within a setting the model still decides for itself how much to think. Two identical questions can land differently. For anyone who has to write a budget, “it depends” is a worse answer than a fixed price, even when it averages out cheaper.

And the advice to re-run your evals sounds more reasonable than it is. Most organisations do not have an eval set. They have a feeling and an invoice. Without your own measurements, “turn it down where quality holds” becomes “turn it down and hope”, and that is the kind of advice you only get to check three months later.

If you take one thing from this piece: build a small set of ten real tasks from your own work, run them at low, medium and high, and look at where it falls apart. That is an afternoon of work, and it answers the question no benchmark answers for you.

Next up: the benchmarks themselves, including the rows where Opus 5 loses.

Dit stuk verscheen ook in het Nederlands: Effort is de knop waar het bij Opus 5 om draait.