I usually do not read model releases for the benchmarks. This time I did, for one reason: the table Anthropic published contains rows where their own new model loses. That is rare in an announcement, and it makes the rest of the table more credible.

Source: Anthropic, Introducing Claude Opus 5.
What genuinely impresses
Three numbers stand out.
On Frontier-Bench v0.1, agentic work in the terminal, Opus 5 goes from 21.1 percent (Opus 4.8) to 43.3 percent. More than double in one generation, and well clear of GPT-5.6 Sol at 34.4. This is the kind of work I do most of, and the jump is big enough to notice outside a test harness.
On ARC-AGI-3, a test built from problems the model has not seen before, it scores 30.2 percent against 1.5 for Opus 4.8 and 7.8 for GPT-5.6 Sol. That is a factor of twenty over the previous version.
On AutomationBench, business tasks carried start to finish, Opus 5 reaches 26.0 percent where the rest of the field sits between 17 and 18. That is the row closest to ordinary office work.
The four rows it loses
And then the more interesting side.
On DeepSWE v1.1, GPT-5.6 Sol wins with 72.7 percent against 68.8 for Opus 5. Not marginally, just better. On FrontierCode v1.1, Fable 5 wins 53.5 to 53.4, which is inside the noise but still means Opus 5 is not the top there. On the Legal Agent Benchmark, Fable 5 wins 13.3 to 11.7. And on HealthBench Professional, Opus 5 sits at 59.8, below both GPT-5.6 Sol (60.5) and Mythos 5 (66.0).
Four of fourteen. For a model positioned as the new default that is honestly reported, and it is a useful signal too: if your work sits heavily in coding agents or in healthcare, “the newest Anthropic model” is not automatically the right answer.
The number that puts it all in perspective
Look at that legal row again. The best model in the table scores 13.3 percent. Opus 5 scores 11.7.
That is not “nearly there”. That is: on this test, the best model available fails in almost nine cases out of ten. If you read the announcement and conclude legal work can now be automated, the evidence that it cannot is in the same table.
The same applies more gently to ARC-AGI-3. Twenty times better than the previous version sounds like a breakthrough, and in relative terms it is. In absolute terms, seventy percent of novel problems remain unsolved.
I do not call that a criticism of the model. I call it the context that falls out of press releases the moment they get retold.
What the table does not tell you
Two things worth knowing.
The benchmarks were replaced, not just filled in. There is no SWE-bench Verified in this table, for years the number everyone watched. There is Frontier-Bench, FrontierCode, DeepSWE and CursorBench. That is partly fair, because old tests saturate once everyone scores ninety percent. But it makes cross-generation comparison hard, and it hands the vendor influence over the yardstick. When the test changes at the same time as the model, you do not know exactly what you are measuring.
It is a vendor measuring its own product. The setup looks careful and the losing rows help, but these are not independent measurements. GDPval-AA and the AA Coding Agent Index come from Artificial Analysis, a third party; most of the rest does not. Wait for independent replication before you put a number in a deck.
What I see myself
I have been working with this model on real jobs since today, so my impression is fresh and no more than that. What stands out is not in the table: it finishes tasks outright more often, and it announces less about what it is going to do and simply does it. That saves turns, and turns are time and money.
Something else stands out too: it stretches its brief. A few times I got back more than I asked for, cleanly done but not requested. That is exactly the behaviour the next part is about, because if you build on the API that is not charm, it is a bug in your prompt.
Dit stuk verscheen ook in het Nederlands: De benchmarks van Opus 5, en de vier rijen waar het verliest.
