
Why Stronger Models Are Often Safer — and Still Not Interchangeable
Anthropic’s Sonnet 5 safety notes show a useful split: mid-tier models can get safer for everyday agent use, while high-risk cyber skills still sit far above them.
When a mid-tier model gets much better at finishing agent jobs, a natural fear follows: does “more capable” also mean “more dangerous”? Anthropic’s Sonnet 5 safety write-up answers with a more nuanced picture. On several everyday safety dimensions, Sonnet 5 looks better than Sonnet 4.6. On high-risk cybersecurity skills, it remains far below flagship-class models — and ships with safeguards that reflect that judgment.
Safer for ordinary agent work
Pre-deployment evaluations, as summarized in the launch post, found Sonnet 5 safer overall than Sonnet 4.6. In agentic safety terms, it is better at refusing malicious requests and resisting hijack attempts in prompt-injection attacks. Hallucination and sycophancy rates are lower. On Anthropic’s automated behavioral audit — which covers a wide range of misaligned behaviors such as cooperating with misuse or deception — Sonnet 5 scores lower, meaning fewer undesirable behaviors.
There is still a gradient. Compared with more capable Opus 4.8 and Claude Mythos Preview, Sonnet 5 still shows somewhat higher rates of misaligned behavior on that audit. That fits Anthropic’s long-running safety ladder story: as base capability rises, alignment outcomes often improve too. Mid-tier sits in the middle — better than its predecessor, not identical to the top of the stack.
Cyber skills are where tiers still diverge sharply
Anthropic is explicit: it did not deliberately train Sonnet 5 on cybersecurity tasks. The model can handle some routine, non-harmful cyber work, but on evaluations of potentially dangerous skills — such as developing software exploits — it performs substantially worse than Opus 4.8 and Mythos 5. In a Mozilla-collaborative test around Firefox 147 vulnerabilities (already patched in Firefox 148), neither Sonnet model produced a full working exploit. Sonnet 5’s partial-success rate rose only slightly over Sonnet 4.6, which Anthropic reads as a side effect of general intelligence gains, not specialized cyber training.
Because Sonnet 5 is still a bit stronger than its predecessor on those tasks, it launches with cyber safeguards enabled by default — the same real-time detection and blocking style used on Opus 4.7 and 4.8, but less strict than the broader blocks described for Fable 5. Anthropic’s stated risk judgment for Sonnet 5 is low overall. For professional cybersecurity work that needs reduced guardrails, the company still points people to Opus 4.8, and notes Sonnet 5’s place in its Cyber Verification Program across Claude’s native platform, AWS, and Microsoft Foundry, with Vertex access following.
How to read the safety ladder as a buyer
- “Safer than the last mid-tier” is not the same claim as “as aligned as the top model”
- Everyday agent usefulness and high-risk cyber capability can move at different speeds
- Default safeguards are part of the product, not only a footnote for security teams
The useful takeaway is to split two questions that marketing often blends. First: will this model refuse the wrong ask and stay steady on ordinary automation? Second: does this tier have the specialized high-risk skills — and the matching access program — that your security workflow actually needs? Sonnet 5’s public safety story says mid-tier can improve on the first without leaping into the second. That is why “stronger and often safer” can be true, while “all strong models are interchangeable” still is not.