
Why Effort Level Matters More Than the Model Name
Anthropic’s Sonnet 5 cost–performance curves make a practical point: you can tune how hard a model tries, instead of only shopping by flagship vs mid-tier.
Model shopping often collapses into a false choice: buy the expensive flagship, or settle for the cheaper mid-tier. Anthropic’s Sonnet 5 launch materials push a different mental model. Alongside the model name, users get an effort dial — spend more compute and tokens when the job is hard, spend less when it is routine.
What the cost–performance curves are saying
In the announcement, Anthropic plots Sonnet 5 against Sonnet 4.6 and Opus 4.8 across effort levels on agentic search (BrowseComp) and computer use (OSWorld-Verified). The plain-language reading: Sonnet 5 is a broad upgrade over the previous Sonnet, and the band of cost–performance options it covers is wider than Opus 4.8’s on these charts. At medium effort, cost efficiency stands out; at the highest effort, some tasks can match Opus 4.8.
That is the practical punchline. You are less forced to pick “cheap” or “good” as permanent identities. You pick a tier for the jobs you run most often, then turn effort up when a messy research task or a stubborn computer-use workflow needs the extra push.
Why higher effort shows up on the bill
Effort is not a free cosmetic setting. Higher effort usually means more tokens, more tool rounds, and longer runs. Anthropic says it raised rate limits across chat, collaboration tools, coding products, and the open platform specifically to accommodate higher-effort usage. Introductory API pricing through 31 August 2026 ($2 / $10 per million input / output tokens, then $3 / $15) also matters when you are stress-testing those dials on real workloads.
One footnote from the same launch is easy to miss: Anthropic later corrected a BrowseComp chart that had used a simpler method and understated Sonnet 5. The update aligned the public chart with the system-card methodology. For readers, the lesson is modest but useful — treat launch charts as living documents, and prefer the method the lab says matches its standard evaluation.
A simple way to choose
If you only remember one workflow from the Sonnet 5 materials, make it this:
- Start from task difficulty, not from brand hierarchy
- Use a mid effort setting as the default for recurring work
- Raise effort only when completion rate or quality clearly needs it
- Keep flagship models for the thin slice of work where the dial is not enough
The deeper product shift is that “which model” and “how hard should it try” are becoming separate knobs. That is good news for budgets, and a warning for naive comparisons: a low-effort mid-tier run and a max-effort mid-tier run are not the same product, even when the model name on the invoice is identical.