Claude Opus 5 costs exactly what its predecessor cost. Five dollars per million input tokens, twenty-five per million out. That sounds like a non-story until you work out what it means: if the price per token is fixed, your bill is decided by the number of tokens. Which is precisely what the effort dial controls.
In this release that is no longer a detail. It is the main story.
What effort is
Effort is a setting with five positions: low, medium, high, xhigh and max. It governs how deeply the model reasons and how many tokens it may spend on a job before answering. The default is high.
You meet it in different places. On the API you set it as a parameter. In Claude Code, xhigh is the default. If you use Claude through the app, the dial is not on your screen, and the model decides per turn how much thinking a question deserves.
What is interesting is not that the dial exists. Opus 4.8 had it too. What is interesting is that the settings mean something different this time.
The surprise sits at the bottom end
The habit until now was: set it high, or the model gets dumb. On Opus 4.8 that was roughly true.
On Opus 5 it is not. The clearest number comes from Zapier’s AutomationBench, which measures whether a model can carry business tasks from start to finish. On its lowest setting, Opus 5 passes more tasks there than any other model does on its highest. Not more than Opus 4.8. More than every other model.
Anthropic says it outright in their own guidance: use low and medium liberally where your evals show quality holds. And: if you carried your effort defaults over from a previous model, run the sweep again. That is unusually direct advice from a vendor that earns per token.
Which setting for what
This is how I divide it up now. Not a law, but a starting point.
| Setting | For | What you notice |
|---|---|---|
low | Summarising, classifying, short questions, anything speed-sensitive | Fast and cheap, at quality that used to need high |
medium | Most daily work: writing, analysis, one-off code questions | The new default if you watch cost |
high | The default value. Precision work with a contained scope | What you were used to, for fewer tokens |
xhigh | Coding and agentic work. The recommended starting point there | Multiple files, longer chains, less giving up |
max | The genuinely hard cases where correctness beats cost | Best result, clearly pricier, occasionally overthinks |
What stands out in that table: there is no longer any reason to sit on high by default without thinking about it. Down is a win, up is a win, standing still is the only option that gains you nothing.
What the numbers say about the bill
Almost every measurement customers have published points the same way: not better answers for more money, but the same or better answers for less.
A financial customer measures nine percentage points higher accuracy across effort levels, with a third fewer turns and tool calls and sixty percent less time. A legal customer reaches comparable performance using twenty-six percent fewer tokens than Opus 4.8 at max reasoning. On CursorBench the model at max lands within half a percent of Fable 5, at half the cost per task.
Translate that into your own situation and you get a different conversation than at previous launches. Not “is the upgrade worth it”, but “where do I have my settings too high”.
The caveat nobody prints alongside
Two things bother me.
Effort makes your costs less predictable, not more. You now have a dial you can steer with, but within a setting the model still decides for itself how much to think. Two identical questions can land differently. For anyone who has to write a budget, “it depends” is a worse answer than a fixed price, even when it averages out cheaper.
And the advice to re-run your evals sounds more reasonable than it is. Most organisations do not have an eval set. They have a feeling and an invoice. Without your own measurements, “turn it down where quality holds” becomes “turn it down and hope”, and that is the kind of advice you only get to check three months later.
If you take one thing from this piece: build a small set of ten real tasks from your own work, run them at low, medium and high, and look at where it falls apart. That is an afternoon of work, and it answers the question no benchmark answers for you.
Next up: the benchmarks themselves, including the rows where Opus 5 loses.
Dit stuk verscheen ook in het Nederlands: Effort is de knop waar het bij Opus 5 om draait.
