i
News
News · 2026-09-22

Anthropic cuts Opus 5.5 pricing as token use stays high

@neuronium_ai @neuronium_ai

Anthropic has released Claude Opus 5.5 at a lower token price while claiming performance on par with Fable 5.1 and GPT-6 Astra in several demanding benchmarks. The model also promises clearer writing, stronger safety controls and new restrictions on distillation attacks. The trade-off is less obvious in practice: Opus 5.5 uses far more output tokens than competing models, so its lower per-token price does not necessarily make each task cheaper.

Cover: Anthropic cuts Opus 5.5 pricing as token use stays high

What shipped

Anthropic set Opus 5.5 pricing at $4 per million input tokens and $20 per million output tokens, down from $5 and $25 for Opus 5. Token prices have fallen by 20%, while cache reads are now 60% cheaper.

The company says total operating costs, including both token prices and token volume, are about 40% lower than with Opus 5. Opus 5.5 uses fewer tokens, and its responses arrive more than 30% faster.

$4per million input tokens
$20per million output tokens
20%token price reduction
40%total cost reduction

For subscribers, 15-minute limits will increase by 20%. Combined with the lower price, Anthropic expects users to complete 25% more work overall. Users will also be able to save a limit reset until they need it.

Anthropic is using coding benchmarks to frame the price-performance comparison:

FrontierCode — Opus 5.5 outperforms OpenAI GPT-6 Astra at roughly 20% of Astra’s cost per task.
Terminal-Bench 4.0 — it matches Astra’s result at 40% of the cost of an Astra task.
CursorBench — it leads GPT-5.6 Sol by 11 points at one-third of Sol’s task cost.

The lower price is also a response to pressure from OpenAI and especially Chinese AI models. Those models trail the performance of more powerful systems but cost only a small fraction as much.

Less Claudish

Anthropic says Opus 5.5 communicates more naturally than earlier models. It should lead with the main point, use less jargon and follow writing-style instructions more precisely. Early testers described its answers as clearer and easier to understand, which Anthropic says makes the model better suited to long-running collaborative work.

That is a direct response to a familiar criticism of Claude: its writing can be formulaic and difficult to follow, a style users sometimes call Claudish.

Screenshot via PAW

Screenshot via PAW

Source: the-decoder.com

Opus 5.5 is the first Opus model with cybersecurity, biology and frontier-language-model-development safeguards at the level of Fable 5.1. When those restrictions are triggered, Anthropic says requests are transparently redirected to another model.

Users can still search for and fix errors in code, but most cybersecurity tasks will be routed to the older Opus 4.8. Requests classified as biology or frontier language model development will go to Opus 5.

Verified organizations can apply to use the model for biological research through the Life Sciences Verification Program. In the coming weeks, Anthropic plans to expand its existing Cyber Verification Program to include Opus 5.5.

The model is available across all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure. Developers using Claude Platform can call it with the identifier claude-opus-5-5.

The safety and control layer

Anthropic plans to apply stricter checks to reinforcement-learning environments. The company considers poorly designed training environments one of the main sources of model behavior that fails to match its intended goals. It is also working on:

better rewards for aligning model behavior;
automatic creation of safety-training scenarios;
stronger security measures and monitoring.

The company is taking a position similar to OpenAI’s recent call for international standards for systems capable of improving themselves. Several researchers have publicly proposed slowing the release of new models until alignment methods catch up with their capabilities.

Anthropic wants a higher safety standard for models that could fully automate AI research, with government policy playing a stronger role in setting that standard. Before release, Opus 5.5 was tested by the external organizations Frontier Design and METR.

Anthropic is also adding defenses against distillation attacks. The company says attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale and build powerful systems without built-in safety mechanisms. Its September 2026 threat report describes illegal distillation activity that Anthropic identified and stopped.

Opus 5.5 includes Preserved Thinking, a feature first introduced in Fable 5.1. It prevents API users from editing Claude’s previous context to extract its reasoning. The restriction applies to Fable 5.1 and Opus 5.5 in API accounts created on or after August 31, 2026.

The model also adds watermarking measures to meet the requirements of the EU AI Act. It can no longer run with Thinking disabled, although it remains available in Zero Data Retention mode.

The benchmark result has a cost

Artificial Analysis gives Opus 5.5 a strong independent showing. At its highest effort level, the model scored 58 on the Artificial Analysis Intelligence Index, the highest result in the index’s history and several points above the previous leader. It led six of the ten evaluations:

On Humanity's Last Exam, it scored 61.4%, versus 59.1% for the previous best result from Fable 5.1.
On SciCode, it scored 66.9%, versus 63.1% for Fable 5.1.
On Terminal-Bench 4.0, it scored 59.6%, matching GPT-6 Astra and beating Opus 5 by 11 points.
Claude Opus 5.5 leads the Artificial Analysis Intelligence Index with 58 points, followed by Claude Fable 5.1 and GPT-6 Astra at 53 each. In the cost-performance chart (bottom), several Opus 5.5 effort levels sit on the Pareto frontier. | Image: Artificial Analysis

Claude Opus 5.5 leads the Artificial Analysis Intelligence Index with 58 points, followed by Claude Fable 5.1 and GPT-6 Astra at 53 each. In the cost-performance chart (bottom), several Opus 5.5 effort levels sit on the Pareto frontier. | Image: Artificial Analysis

Source: the-decoder.com

According to Artificial Analysis, Opus 5.5 puts Anthropic level with GPT-6 Astra on Terminal-Bench 4.0 and AutomationBench-AA, while widening its lead in agentic knowledge work. On the closed AA-Briefcase benchmark, it achieved an Elo rating of 1,822, 143 points ahead of Fable 5.1.

This is also the first time an Anthropic model has beaten GPT-5.6 Sol on presentation quality. In GDPval-AA, OpenAI’s benchmark for office work, Opus 5.5 beat Astra in every reasoning mode except low.

On GDPval-AA v2.1, Claude Opus 5.5 achieves higher Elo scores than competitors at comparable or lower cost per task across nearly all effort levels. Only at the "low" setting does Astra come out ahead. | Image: Anthropic

On GDPval-AA v2.1, Claude Opus 5.5 achieves higher Elo scores than competitors at comparable or lower cost per task across nearly all effort levels. Only at the "low" setting does Astra come out ahead. | Image: Anthropic

Source: the-decoder.com

The efficiency trade-off is substantial. At maximum effort, Opus 5.5 uses about 119,000 output tokens per task, compared with 73,000 for Opus 5, 78,000 for Fable 5.1 and 27,000 for GPT-6 Astra. Its lower token price keeps the cost of a task roughly at Opus 5’s level, but does not make the task cheaper.

My read is that Opus 5.5 is a pricing and positioning move as much as a model upgrade. Anthropic can claim parity or superiority on important evaluations while making the API look more competitive, but the headline price hides a resource-intensive reasoning profile. Four of its five effort modes sit on the Pareto frontier: among models scoring above 50 on the index, they match or beat all others on task cost. That is a strong result, but it is not the same as making advanced reasoning inexpensive.

The quieter question is what customers will pay for clearer collaboration and better benchmark scores once long tasks consume six figures of output tokens. Anthropic has narrowed the price gap without removing the underlying tension between model capability, safety controls and the amount of computation each answer requires.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X