A model built for frequent, narrow tasks
Anthropic’s example from financial AI company Rogo shows the intended division of labor: a larger model prepares a presentation, then Haiku finds a revenue figure for one of its slides.
Alex Wang, an applied AI specialist at Rogo, said Haiku was accurate enough for that work and fast and inexpensive enough to run often.
The pricing favors short requests. Below 100,000 tokens, Haiku 5.5 costs 90% less than Haiku 4.5. Above that threshold, it costs $0.50 per million input tokens and $2.50 per million output tokens—50% less than Haiku 4.5. Anthropic has not specified how requests exactly 100,000 tokens long are billed or which tokens count toward the threshold.
Anthropic estimates that about 90% of requests to Haiku 4.5 fall into the short-request category. It puts average savings on typical workloads at about 75%, accounting for an updated tokenizer that uses slightly more tokens for equivalent work.
Price parity does not settle the comparison
At the lower price tier, Haiku matches four GPT-6 Luna rates, including cache reads. But matching prices does not make either model the cheapest on the market, or establish that they do the same work.
Anthropic’s comparison also includes higher rates for input volumes above 272,000 tokens. The standard promotional prices are in effect through December 31, 2026. The figures exclude caching, batch-processing discounts, tools, regional surcharges and individually negotiated rates. They compare cost, not model capabilities.
The thresholds matter: Haiku’s upper tier is five times its lower one, while Luna’s surcharge begins after 272,000 input tokens. And token prices alone do not determine the cost of finishing a task; token consumption, retries and accuracy all matter.
Anthropic’s benchmark table shows Haiku 5.5 ahead of Haiku 4.5 and GPT-6 Luna on the listed tests, with Sonnet 5.5 scoring higher. The results come from Anthropic and have not been independently verified. OSWorld 2.1 used an offline sample; Sonnet’s FrontierCode result used high effort, and a dash means no result was provided. The first two rows report scores, not percentages.
Haiku 5.5 also adds configurable effort levels, with medium selected by default. On Anthropic’s Terminal-Bench chart, Haiku scores about 39% at maximum effort and about 20% at medium. That gap makes the headline benchmark number less useful without knowing how much effort—and therefore what operating cost—produced it.
What the launch leaves unanswered
I think the more consequential question is not whether Haiku’s low-end rate matches Luna’s, but how reliably it completes narrow tasks when called repeatedly. Anthropic’s own positioning suggests a useful boundary: let Haiku handle frequent, bounded work and leave harder tasks to larger models. The benchmark results support that division, but they do not establish how often a real agent will need retries or escalation.
Customer reports offer examples, not a common performance measure. Box’s vice president of AI products, Yashodha Bhavnani, reported an 11-point improvement over Haiku 4.5 and roughly half the latency, without specifying the scoring scale. Asana lead software engineer Aaron Vinh said task latency fell by more than 30% and per-turn agent inference became up to 2.5 times faster than with an unnamed model. HubSpot lead software engineer Ziv Klapow reported an average score of 92.8% across three runs of the company’s internal CRM evaluation, the best result among the models it tested. These company-provided accounts do not allow direct comparisons of throughput across vendors.
Anthropic calls Haiku its fastest model at standard speed, while acknowledging that Opus in Fast Mode is faster. The materials reviewed for this article do not state a tokens-per-second figure. Without throughput data, the lower price and customer latency reports leave a practical cost question open: how much work can the model finish per unit of time?
More routes to adoption, with limits
Anthropic has also lowered Sonnet 5.5 cache-read pricing from $0.20 to $0.10 per million tokens. It estimates savings of about 20% on typical agent workloads, depending on how often an application reuses cached context.
Monthly API credits begin going out this week: $100 for Max 5x subscribers, $200 for Max 20x subscribers, and up to $500 across Team subscribers. The credits can be used on any model through Anthropic’s platform. The draft materials do not explain how Team credits are allocated or whether they carry over.
Haiku is available through Anthropic’s platform, AWS, Google Cloud and Microsoft Azure, under the API identifier claude-haiku-5-5. Updated Python and TypeScript packages add beta features for controlling browsers and computers. Some existing Azure and Google Cloud customers will receive the lower Sonnet cache price over the next few days.
Compared with Haiku 4.5, the new model has stricter cybersecurity limits: standard safeguards block penetration testing. Organizations seeking broader cybersecurity or biology capabilities can apply through Anthropic’s verification programs.
The launch makes frequent use easier to justify on price, but leaves task-level economics unproven. If Haiku can reliably take the narrow work off larger models, its value will be measured not by the rate card alone, but by how often that handoff actually holds.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X