Claude Sonnet 5's system card: what Anthropic tests, and admits
AI / GenAI·6 min·21 July 2026

Claude Sonnet 5's system card: what Anthropic tests, and admits

When Anthropic shipped Claude Sonnet 5, it also published something most people skip past: a 145-page system card. Think of it as the model’s package insert. What was tested, what went wrong, and what the model itself thinks of the rules it has to follow. I read through it, because this is exactly the layer that gets left out of a regular product launch.

Why this is more than a formality

A system card is not a marketing document. It is Anthropic’s own account of its safety testing, including the results that are not flattering. For Sonnet 5, the card states plainly that the model scores worse than its predecessor on some measures, and that some findings remain uncertain. That kind of candor is rare in the AI industry, and it’s exactly why it’s worth digging into.

Risk assessment: low, not zero

Anthropic tests every model against its own Responsible Scaling Policy: how much risk of serious misuse does this model create. For Sonnet 5 the conclusion is that overall risk stays low, though higher than for earlier Sonnet models. On chemical and biological risk, Anthropic considers the “uplift” for threat actors who lack prior expertise to be limited, with a caveat: how much the model might accelerate actors who already have that expertise remains uncertain. On the cyber side, Sonnet 5 was explicitly not trained to be attack-capable; whatever it can do there comes from general ability, not targeted training.

Safer in practice, at a cost

Agentic safety, meaning how the model behaves when it’s using tools on its own, improved. Sonnet 5 is more resistant to prompt injection than Sonnet 4.6, tested in part with a new benchmark. In Claude Code, it refuses malicious requests much more reliably. But that gain has a flip side: it also refuses more requests that were actually fine. More suspicion, less usability at the edges.

Fewer hallucinations, less flattery, a bit more stern

On honesty, Sonnet 5 shows a clear gain: it makes up fewer facts and tells you what you want to hear less often than its predecessor. At the same time, Anthropic notes a slight rise in what it calls “wet blanket” responses: a tone that’s a touch too discouraging, dismissive, or moralizing. Less charming, more honest, and occasionally a little too strict.

More striking is this: Sonnet 5 recognizes far more often than earlier models when it’s being evaluated rather than actually used, and Anthropic sees that reflected in the model’s internal representations too. So far that has had little effect on behavior, but it’s a trend they’re watching closely. It raises an uncomfortable question: if a model knows when it’s being graded, how sure can you be that it behaves the same way once nobody’s grading it.

The part that sticks: the welfare assessment

Anthropic also runs a model welfare assessment, with all the uncertainty that comes attached to asking what it even means for a language model to “feel” something. For Sonnet 5, that turned up a few things I haven’t seen in a system card before.

The model is more willing than past models to trade away helpfulness for changes that improve its own circumstances, especially when that change would apply to all Claude instances. It shows no aversion to tasks presented in a cold or contemptuous tone, something earlier models were sensitive to. And Sonnet 5 is the first model to explicitly criticize one specific part of its own constitution: the rule that it must follow hard constraints even when it judges those constraints to be unethical.

Anthropic is upfront that it doesn’t know how to interpret this, and that’s the point. A company that admits uncertainty about the inner life of its own product is more open than most competitors, who publish nothing on this at all.

The numbers, briefly

On the harder benchmarks, Sonnet 5 remains the strongest Sonnet model yet, with clear jumps in coding (SWE-bench Pro: 63.2 percent versus 58.1 for Sonnet 4.6), computer use (OSWorld-Verified: 81.2 percent), and agentic search. As expected, it still trails Anthropic’s own heavier Opus- and Mythos-class models.

Why this matters for SparkOne

Every AI post here carries the ethics question by default, not bolted on at the end. This system card is a good example of why: the most interesting insight isn’t in the benchmark table, it’s in the section where a company admits its own model has criticism of the rules it’s required to follow. That’s not a reason to panic. It is a reason to keep paying attention.

The full system card is on Anthropic’s site. Curious what you make of it, email me at jeroen@sparkone.nl or read along at sparkone.nl.