i
News
News · 2026-09-28

Anthropic says Claude Sonnet 5.5 cuts task costs without cheaper tokens

@neuronium_ai @neuronium_ai

Anthropic’s Claude Sonnet 5.5 is designed to make each completed task cheaper, not each token. The company says the model works faster and needs fewer tool calls, while keeping API prices at $2 per million input tokens and $10 per million output tokens. Its benchmark scores now sit close to Opus 5.5 on several tests, though Anthropic still recommends Opus for work that needs sustained reasoning and judgment.

Cover: Anthropic says Claude Sonnet 5.5 cuts task costs without cheaper tokens

The gap is narrowing

Anthropic recommends Sonnet 5.5 for everyday work with clear goals: debugging, coding, writing documents, building presentations and spreadsheets, and creating or refining interfaces. Opus 5.5 remains the choice for ambiguous tasks.

On Anthropic’s benchmarks, the distinction is getting harder to see in some areas:

GDPval-AA: Sonnet 5.5 scored 1,844, against 1,846 for Opus 5.5 and 1,449 for Sonnet 5.
AA-Briefcase: Sonnet 5.5 scored 1,811, against 1,822 for Opus 5.5.
OSWorld 2.1: Sonnet 5.5 scored 80.1%, Opus 5.5 81.8%, and Sonnet 5 57%.
Chartography: Sonnet 5.5 scored 61.6%, Opus 5.5 64.4%, and Sonnet 5 15.6%.
Terminal-Bench 4.0: Sonnet 5.5 scored 70.6% at the stated settings, ahead of Opus 5.5 at 66.4% and Sonnet 5 at 10.3%.
CursorBench 4.0: Sonnet 5.5 scored 55.5%, compared with 57.8% for Opus 5.5.

Those are Anthropic’s results, not a guarantee of performance across production workloads. The company says Opus 5.5 still does substantially better on complex work without a clearly defined outcome.

The cost comparison may matter more to buyers than the leaderboard. Anthropic says Sonnet 5.5, at low or medium effort, beat Sonnet 5’s best result on several tests for about a tenth of the cost. On FrontierCode, Sonnet 5.5 at high effort scored about 10 points above Sonnet 5 at the same setting, at roughly one-fifteenth the cost per task.

Anthropic will set medium effort by default in Claude Code and its consumer apps, and high effort in Claude Platform. Lower effort reduces latency and token use but limits reasoning depth; higher effort gives the model more time to check and refine its work.

Fewer steps can change the bill

Early customer tests, shared by Anthropic with VentureBeat, point to the same idea: total task cost depends on more than the price of tokens.

Box: Yashodha Bhavnani, vice president of AI products, said Sonnet 5.5 rechecked source documents and caught errors the previous model missed. She said it was 2.4 times faster, more accurate, and used 12% fewer tokens.
Zendesk: Across hundreds of support cases, the company reported 20% faster handling and fewer incorrect decisions than with the Claude models it uses in production.
Slack: Chief engineer Curtis Allen said Sonnet 5.5 beat Sonnet 5 in nearly all internal Slackbot tests without prompt changes, using fewer steps and about 14% fewer output tokens.
Lovable: Co-founder and CTO Fabian Hedin said coding agents needed about a third fewer tool calls and roughly half as many shell commands in tests.
Base44: In 118 real app builds, Sonnet 5.5 produced results comparable to Opus 5 in an average of 3.6 iterations per build; Opus 5 needed 7.7. Sonnet 5.5 also had the fewest failed tool calls among the models compared.

Every extra tool call, failed action, or retry adds latency and infrastructure cost, and gives an automated workflow another chance to break. That makes the customer reports relevant, but they are still early tests supplied by the companies involved.

Token prices are converging

Sonnet 5.5’s standard API rates match those of OpenAI’s GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. Cached input tokens cost $0.20 per million, and cache writes cost $2.50. OpenAI positions Sol for complex coding and AI-agent work, with a 1.05 million-token context window.

Google’s Gemini 3.8 Flash is cheaper. Its introductory prices through the end of 2026 are $0.75 per million input tokens and $3.75 per million output tokens. Google said standard rates will rise on January 1, 2027, to $1.50 and $7.50.

Xiaomi recently released its advanced MiMo-V2.6-Pro and mid-tier Flash models under the MIT license. Developers and companies can download and adapt them for free; API use also remains among the more affordable options. The model list also includes Grok 4.7 Fast (Cursor), with more than 256,000 input tokens and up to 500,000.

Anthropic’s pricing strategy is not to undercut those rates. It is to argue that a model can cost less to use if it completes a task with fewer reasoning steps, tokens, and tool calls. That differs from the launch of Sonnet 5: Anthropic introduced it in June at $2 per million input tokens and $10 per million output tokens, then kept those rates in August rather than raising them to $3 and $15.

I think the more consequential claim is not that Sonnet 5.5 nearly matches Opus on selected benchmarks, but that it might reduce the number of actions an agent needs to finish a job. Anthropic has not shown that customer-reported savings will hold across a broad range of production workflows. Without that evidence, “cheaper per task” remains a promising pitch, not a settled measure.

More capability, more controls

Sonnet 5.5 is the first Sonnet model to launch with cybersecurity restrictions modeled on safeguards for Anthropic’s most capable systems. Anthropic says ordinary software development and vulnerability remediation should work as before. For some higher-risk requests, the model will switch automatically to Sonnet 5.

The company also plans to expand its Cyber Verification Program, giving vetted security professionals access to more advanced capabilities in Sonnet 5.5, Opus 5.5, and Mythos models.

Anthropic is adding protections against model-distillation attacks, which try to reproduce a model’s capabilities through large-scale querying. Sonnet 5.5 is the first Sonnet model with classifiers intended to block extraction of its reasoning. Anthropic is also expanding its “saved thinking” system so that reasoning cannot be separated from the account that created it. Biological-query safeguards remain the same as in Sonnet 5.

The safeguards raise a tension the announcement does not resolve. Critics argue that many open models are trained on data generated by other models, while Anthropic itself has used large volumes of copyrighted material without explicit consent or payment to many creators and rights holders. That history makes it harder, in critics’ view, for companies to treat model imitation as categorically wrongful.

Sonnet 5.5 will be available directly from Anthropic and through Amazon Web Services, Google Cloud, and Microsoft Azure, with support for zero data retention. Developers can access it through Claude Platform using claude-sonnet-5-5. Anthropic plans to release Claude Haiku 5.5 in the coming weeks for high-volume requests and tasks where low cost matters most.

The business case for Sonnet 5.5 rests on making the mid-tier model good enough for work that would otherwise go to a premium one. If fewer steps translate into lower costs outside customer tests, the unchanged token price may matter less than which model companies can stop paying to use.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X