i
DATAIST
News · 2026-08-31

Claude Sonnet 5 nears Opus 4.8 at $2 per million input tokens

@neuronium_ai @neuronium_ai

Claude Sonnet 5 is available across every plan at $2 per million input tokens and $10 per million output tokens, and the claim attached to it is that its performance sits close to Opus 4.8 for a lower price. It is the default model on the Free and Pro plans, available to Max, Team and Enterprise users, and exposed to developers as claude-sonnet-5. Anthropic had planned to move it to standard rates of $3 and $15 on September 1; on August 10 the company confirmed the launch price is permanent instead.

Cover: Claude Sonnet 5 nears Opus 4.8 at $2 per million input tokens

Claude Sonnet 5 is available across every plan at $2 per million input tokens and $10 per million output tokens, and the claim attached to it is that its performance sits close to Opus 4.8 for a lower price. It is the default model on the Free and Pro plans, available to Max, Team and Enterprise users, and exposed to developers as claude-sonnet-5. Anthropic had planned to move it to standard rates of $3 and $15 on September 1; on August 10 the company confirmed the launch price is permanent instead.

For a lot of developers the agent era started with Sonnet. Claude Sonnet 3.5, 3.6 and 3.7 were the first models to show real ability at coding and tool use. Since then the clearest agentic progress has come from the Opus line, and Sonnet has been the cheaper seat rather than the capable one. Sonnet 5 is the attempt to narrow that, with gains over Sonnet 4.6 across reasoning, tool use, coding and knowledge work.

The comparison Anthropic publishes is not model against model but curve against curve. Charts plot Sonnet 5, Sonnet 4.6 and Opus 4.8 at different effort levels on BrowseComp, an agentic search evaluation, and OSWorld-Verified, a computer-use benchmark. Sonnet 5 improves on Sonnet 4.6 across the whole range and covers more of the cost-performance space than Opus 4.8 does. At medium effort it is markedly cheaper for the work done; at high effort it reaches Opus 4.8 results on some tasks. Effort is a dial the user sets, on both models.

Cost-performance curves at different effort levels. The previous best Sonnet model (Sonnet 4.6) lagged well behind Opus 4.8. Sonnet 5 offers a wider range of cost-performance tradeoffs than Sonnet 4.6, and in some cases matches Opus 4.8 capability levels.

Cost-performance curves at different effort levels. The previous best Sonnet model (Sonnet 4.6) lagged well behind Opus 4.8. Sonnet 5 offers a wider range of cost-performance tradeoffs than Sonnet 4.6, and in some cases matches Opus 4.8 capability levels.

Source: anthropic.com

Separate figures show how throughput and input and output token costs shift for Sonnet 5 and Opus 4.8 as the effort level changes.

Source: anthropic.com

Early-access partners reported that Sonnet 5 is noticeably better at agentic work than previous versions: it finishes complex tasks where earlier Sonnets stalled, and checks its own output without being asked to.

Pre-release safety testing also came out ahead of Sonnet 4.6. In agentic settings the model is better at refusing harmful requests and at resisting attempts to take control of it through prompt injection, and its rates of hallucination and sycophancy are lower. In an automated behavioral audit covering a very wide range of deviations, including assisting abuse and deception, Sonnet 5 scored as generally safer than Sonnet 4.6 — but showed somewhat more undesirable behavior than the more capable Opus 4.8 and Claude Mythos Preview.

Misaligned behavior rates in an automated behavioral audit that probes an extremely wide spectrum of undesirable behavior across a broad range of situations and contexts (for the full list and per-behavior results, see section 6.4 of the Sonnet 5 system card). Sonnet 5 shows generally lower rates of misaligned behavior than

Misaligned behavior rates in an automated behavioral audit that probes an extremely wide spectrum of undesirable behavior across a broad range of situations and contexts (for the full list and per-behavior results, see section 6.4 of the Sonnet 5 system card). Sonnet 5 shows generally lower rates of misaligned behavior than

Source: anthropic.com

On cybersecurity the model was deliberately not trained. It handles some ordinary, harmless cyber tasks and is distinctly worse than Opus 4.8 and Mythos 5 at the potentially dangerous ones, such as writing malware to exploit vulnerabilities. In one evaluation, models were asked to build exploits against the Firefox browser; Sonnet 5 never produced a fully working one, though it reached partial success slightly more often than Sonnet 4.6. Anthropic attributes that shift to the model getting generally smarter rather than to any targeted training.

Because it does edge past its predecessor there, Sonnet 5 ships with cyber protections enabled by default, detecting and blocking dangerous use in real time, matching the protections on Claude Opus 4.7 and 4.8. Anthropic considers its overall cyber risk low, so those protections are looser than Fable 5's, which blocks a substantially wider range of security tasks. Sonnet 5 is also covered by the cybersecurity verification program on Claude Platform, Claude Platform on AWS and Claude in Microsoft Foundry hosted on Azure and Anthropic, with Claude in Google Vertex to follow; organizations already enrolled do not need to reapply. For security work under lighter restrictions the recommendation is Claude Opus 4.8.

Here is where the pricing story gets less generous than it reads. The launch price held at $2 and $10 instead of rising to $3 and $15, which is a real concession. But Sonnet 5's updated tokenizer turns the same input into roughly 1.0 to 1.35 times as many tokens, depending on content type — the same change made earlier in Claude Opus 4.7. Price per token is flat; tokens per document are not. Anthropic states both numbers and never puts them in the same sentence, and the one that decides your bill is the product of the two. Rate limits were raised in Chat, Cowork, Claude Code and Claude Platform precisely because high effort levels consume more tokens, which is the company conceding the point in operational form.

The second thing worth naming: the model with more misaligned behavior in the audit than Opus 4.8 is the one now running by default for everyone on Free and Pro. That is defensible — it is still safer than the Sonnet it replaces, and volume has to go somewhere cheap — but the safety ranking and the distribution ranking point in opposite directions, and the announcement presents the first without mentioning the second. To Anthropic's credit, it also went back on June 30 and corrected a BrowseComp chart whose simpler methodology had understated Sonnet 5; the revised version follows the system card's approach, with a 10 million token budget, context compaction and programmatic tool calling. Publishing a correction that flatters your own model is a strange kind of honesty, but it is honesty.

Several other things landed alongside the model. Anthropic is opening a research version of the Model Hardware Standard, a shared specification for AI agents operating physical devices safely, initially to a set of research labs and advanced manufacturers. From today, 10,000 scientists worldwide can use Claude free: verified research group leaders qualify for the Claude Team plan and can add colleagues on standard seats at no cost, or on higher-tier seats at $15 a month, for up to a year. A $5 million grant program is launching for independent research into how AI affects user wellbeing. Earlier, on April 26, Sonnet and Haiku limits were raised at every usage tier, with Claude Platform keeping its three tiers: Start, Build and Scale. The full evaluation record is in the Claude Sonnet 5 system card.

What Anthropic has actually shipped is not a cheaper model so much as a dial that makes the Sonnet-versus-Opus question negotiable per request. That is better for buyers who know what their workload needs, and worse for everyone else: the same model name now covers a span of cost and capability wide enough that "we run Sonnet 5" stops describing anything in particular.