Today, Anthropic released Claude Haiku 5.5, which the company says is its cheapest, fastest, and most capable small model to date.

Claude Haiku 5.5

Haiku 5.5's average running cost is about 75% lower than that of the previous generation, Haiku 4.5.

And while it has become so much cheaper, its scores in coding, computer use, and knowledge work are all a big step up from Haiku 4.5.

Anthropic also announced two other things this time: Sonnet 5.5's cache-read price has been cut in half, and Claude Max and Team subscribers will receive an additional amount of API credits each month.

Benchmarks

Let's first look at the official report card.

Haiku 5.5 benchmarks
Claude Haiku 5.5 benchmark table comparing Haiku 5.5, Haiku 4.5, GPT-6 Luna, and Sonnet 5.5 across knowledge work, computer use, multidisciplinary reasoning, agentic coding, and visual reasoning. Sonnet 5.5 is shown for reference.

Besides the previous generation, Haiku 4.5, the table also brings in GPT-6 Luna for a beating, while Sonnet 5.5 on the far right is there for reference only.

The biggest improvement is in computer use. On OSWorld 2.1 (offline subset), Haiku 5.5 scored 72.4%, while Haiku 4.5 scored just 15.7%, and GPT-6 Luna scored 48.9%.

In agentic coding, Haiku 4.5 essentially turned in a blank paper on Terminal-Bench 4.0, scoring 0.0%... But Haiku 5.5 reached 39.2%, and GPT-6 Luna scored 16.4%.

On Humanity's Last Exam, Haiku 5.5 scored 45.9% without tools and 57.4% with tools, while the previous generation scored just 10.2% and 18.7%, respectively.

On the two knowledge-work benchmarks, GDPval-AA v2.1 and AA-Briefcase v1.1, Haiku 5.5 scored 1620 and 1578, respectively, both more than twice Haiku 4.5's scores of 735 and 614.

On the visual-reasoning benchmark Chartography, its score also rose from 6.4% to 46.4%.

Of course, there is still a gap between it and Sonnet 5.5. On Terminal-Bench 4.0, Sonnet 5.5 scored 70.6%, more than 30 points higher than Haiku 5.5. Complex agentic coding tasks should still be handed to Sonnet 5.5 and Opus 5.5.

Haiku 5.5 is suited to narrower tasks, such as context compaction, summarization, and serving as a subagent. Running these tasks with Claude used to be simply not cost-effective.

Adjustable Effort

Haiku 5.5 is the first Haiku model with an effort setting, with five levels in all: Low, Med, High, Xhigh, and Max.

As with Claude's other models, you can now decide for yourself whether each task calls for saving money or more intelligence.

The company released three "score / cost" curves showing the different effort levels.

OSWorld performance at each effort level
Agentic computer-use performance by effort level on the OSWorld 2.1 offline subset. The chart compares partial-credit scores against cost per attempt in USD on a logarithmic scale for Haiku 5.5, Sonnet 5.5, Haiku 4.5, and GPT-6 Luna, with Haiku 5.5's Low, Med, High, Xhigh, and Max levels labeled.

On OSWorld, Haiku 5.5 reaches 72.4% when set to Max, already higher than Sonnet 5.5's two lowest levels, while costing a little less per task.

GDPval-AA performance at each effort level
Real-world knowledge-task performance by effort level on GDPval-AA v2.1. The chart compares reported Elo against cost per task in USD on a logarithmic scale for Haiku 5.5, Sonnet 5.5, Haiku 4.5, and GPT-6 Luna, with Haiku 5.5's Low, Med, High, Xhigh, and Max levels labeled.

Things do not look quite as good on GDPval-AA. At the cheaper levels, GPT-6 Luna scores slightly higher for the same money. Haiku 5.5 relies on Xhigh and Max to push its score up.

HLE performance at each effort level
Multidisciplinary-reasoning performance by effort level on Humanity's Last Exam without tools. The chart compares scores against cost per attempt in USD on a logarithmic scale for Haiku 5.5, Haiku 4.5, and Sonnet 5.5, with Haiku 5.5's Low, Med, High, Xhigh, and Max levels labeled.

On Humanity's Last Exam, increasing the effort from Xhigh to Max does not move the score up much further.

Pricing

Pricing table
Pricing table in USD per 1 million tokens for Haiku 5.5 prompts up to and over 100,000 tokens, Haiku 4.5, and Sonnet 5.5. Rows list cache reads, cache writes, input tokens, and output tokens.

Haiku 5.5's pricing is split into two tiers based on prompt length, with 100,000 tokens as the dividing line.

Price per 1 million tokens
English translation of the source's price-per-million-token table. For prompts up to 100,000 tokens, Haiku 5.5 input costs $0.10, output $0.50, cache reads $0.01, and cache writes $0.13. For prompts over 100,000 tokens, the prices are $0.50, $2.50, $0.05, and $0.63. Haiku 4.5 prices are $1.00, $5.00, $0.10, and $1.25, respectively.

That works out to 90% cheaper for requests within 100,000 tokens, and 50% cheaper for requests over 100,000 tokens.

On Haiku 4.5, about 90% of requests fall within 100,000 tokens.

So why does the average come out to only 75%?

The company explains in a footnote that this average also factors in changes in token usage. Haiku 5.5 has switched to a new tokenizer similar to those in Sonnet 5.5 and Opus 5.5, which uses a few more tokens to do the same thing.

Compare that with Sonnet 5.5, whose input costs $2.00 and output costs $10.00.

For Haiku 5.5 requests within 100,000 tokens, both the input and output prices are just one-twentieth of Sonnet 5.5's.

Use Cases

Haiku 5.5 is also currently Claude's fastest model. The footnote does note, however, that this comparison uses each model's standard speed; Opus is still faster with Fast Mode enabled.

The company has lined up a few main categories of work for it.

  • High-volume, repetitive tasks such as summarization, context compaction, database queries, and classification
  • Serving as a subagent for Opus 5.5 and Sonnet 5.5 during coding
  • Speed-sensitive scenarios such as real-time customer service and browser operations

Several companies that received early access to the model also shared their own figures.

Asana ran it through the evaluation set for its agent product AI Teammates. Compared with the model they currently use, task-completion latency dropped by more than 30%, and single-turn inference was up to 2.5 times as fast.

HubSpot tested the model in a simulated CRM environment. Across three runs, Haiku 5.5 averaged 92.8%, the highest score ever seen on that evaluation.

AlphaSense's document question-answering feature makes about 8 million calls a week in production. They ran 400 queries, and Haiku 5.5 scored 0.84, while Haiku 4.5 scored 0.76.

In Box's early tests, Haiku 5.5 scored 11 points higher than Haiku 4.5, with latency at about half as much.

Cognition added it to Devin Fusion's sidekick lineup. With Haiku 5.5 as its sidekick, Fusion can still maintain a FrontierCode score of 66.2, while bringing down both cost and latency.

This combination is now available in Devin CLI, with Opus 5.5 taking the lead.

Rogo's Alex Wang said:

While the larger model is working on a deck, a Haiku 5.5 subagent dives into the 10-K and finds the line of segment revenue the deck needs. It is accurate enough that we trust it with this work, and fast enough and cheap enough that we can run it at scale.

Safety

On alignment evaluations, Haiku 5.5 improves on Haiku 4.5 almost across the board, with much less misbehavior and less willingness to cooperate with abuse.

In terms of safeguards, it is stricter on cybersecurity than Haiku 4.5, but somewhat less restrictive than other recent models.

Compared with Sonnet 5.5, it allows more defensive tasks, while still blocking techniques such as penetration testing that are more likely to be used by attackers.

Its biological safeguards are consistent with those of Sonnet 5, Sonnet 5.5, and Opus 5.

Sonnet Price Cut and API Credits

Sonnet 5.5's cache-read price has dropped from $0.20 per million tokens to $0.10.

Cache reads account for a large share of the tokens consumed in agent tasks, so after the price cut, the cost of running most agent tasks with Sonnet 5.5 will drop by about 20%.

Exclusive to subscribers, all Max and Team subscriptions will receive an additional amount of API credits each month starting this week, which can be used on Claude Platform.

Max 5x receives $100 per month, Max 20x receives $200 per month, and Team receives up to $500, shared among team members.

These credits can be used with any model. Anthropic hopes people will use them to try building tools, applications, and agents that call the API.

Haiku 5.5 is now available on all platforms, including AWS, Google Cloud, and Microsoft Azure.

The model ID on Claude Platform is claude-haiku-5-5.

Official announcement: https://www.anthropic.com/claude-haiku-5-5

Official tweet: https://x.com/claudeai/status/2107894039626277339

Sources and translation

Anthropic Haiku 5.5 announcement ↗

Complete English translation of the supplied article. Original author not identified in the PDF.

Benchmark scores are provider-reported and have not been independently reproduced by SaveMyToken. The linked X post and full system-card evaluation settings were not independently verified.