Anthropic Releases Claude Haiku 5.5 - One-Tenth of Haiku 4.5's Rates for Prompts Up to 100K Tokens, and Sonnet 5.5 Cache Reads Cut in Half
Anthropic released Claude Haiku 5.5 on October 7, 2026. At $0.10 input and $0.50 output (prompts up to 100K tokens), its rates are one-tenth of Haiku 4.5's, and the company estimates it costs about 75% less on average. Code written for Haiku 4.5 may need changes, as budget_tokens now returns a 400 error. Sonnet 5.5 cache reads also dropped to $0.10.
On October 7, 2026, Anthropic released Claude Haiku 5.5 (claude-haiku-5-5), its new small model1. The company pitches it for large volumes of repetitive work such as summarization, classification, and database lookups, and as a coding subagent working under Opus 5.5 or Sonnet 5.5. Anthropic calls it its lowest-cost, quickest, and strongest small model to date.
It is the third model in the 5.5 family, after Opus 5.5 on September 22 and Sonnet 5.5 on September 28; both of those launches said Haiku 5.5 would follow within weeks. The same announcement also lowered the price Sonnet 5.5 charges for cache reads.
Two price tiers, split by prompt length
Here is the published price table, in US dollars per million tokens1. Haiku 5.5 is the only model in it whose rates depend on whether the prompt is at most 100K tokens or longer.
| Item | Haiku 5.5 (up to / over 100K) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Input | $0.10 / $0.50 | $1.00 | $2.00 |
| Output | $0.50 / $2.50 | $5.00 | $10.00 |
| Cache reads | $0.01 / $0.05 | $0.10 | $0.10 |
| Cache writes | $0.125 / $0.625 | $1.25 | $2.50 |
For prompts up to 100K tokens, every rate in the table is exactly one-tenth of Haiku 4.5’s; above that threshold, each is half. Comparing Haiku 5.5’s own two tiers, crossing the line multiplies the rate by five. Workloads that pass whole long documents in a single request need to keep that boundary in mind.
The $0.10 input and $0.50 output rates match the prices of GPT-6 Luna, released on September 22. Looking only at the up-to-100K tier, the two companies now carry the same price tag.
”About 75% cheaper” is the company’s average estimate
Anthropic says that on average Haiku 5.5 costs roughly 75% less to run than Haiku 4.51. That is a smaller figure than the 90% cut in unit prices, and a footnote spells out the assumptions behind it.
According to the footnote, 90% of requests to Haiku 4.5 were 100K tokens or shorter. The rest fall into the tier where the discount is only 50%. On top of that, Haiku 5.5 has a new tokenizer, so a given task consumes slightly more tokens. The 75% figure is Anthropic’s own average with both factors built in, not a number checked by a third party.
The migration guide is more concrete about the token increase. Haiku 5.5 shares the newer tokenizer used by Claude 4.7 and later models, and identical input text is counted as roughly 30% more tokens than on Haiku 4.5, with the actual difference varying by content3. This is the same tokenizer change covered in our late-August piece on the Sonnet 5 price freeze. Because the way tokens are counted changes, there is no way to know where your bill lands until you recount representative prompts against claude-haiku-5-5. The migration guide likewise asks developers to recount rather than reuse token counts measured on Haiku 4.5.
Swapping the model name is not enough
This is where switching from Haiku 4.5 needs the most care. The release notes state plainly that code built for Haiku 4.5 may fail once pointed at Haiku 5.52. The items listed are that manual extended thinking (budget_tokens) now gets a 400 error, that adaptive thinking is enabled by default so a reply may open with thinking blocks, and that the same text is counted as more tokens.
The main request patterns the migration guide says will fail with a 400 error are3:
thinking: {"type": "enabled", "budget_tokens": N}is rejected. Switch to{"type": "adaptive"}and set thinking depth throughoutput_config.effort- Drop
temperature,top_p, andtop_k. Anytop_kvalue is rejected, as is atemperatureother than 1 or atop_pother than 0.99 - A prefill, where
messagesends with an assistant turn, is rejected even when thinking is turned off - Computer use integrations that declare
computer_20250124are rejected when called via the Claude API or Google Cloud, and have to move tocomputer_toolset_20260801
Some behavior changes without producing an error. Thinking tokens count against max_tokens, so a small limit tuned for Haiku 4.5 can end a response after a thinking block before any answer text appears. Code that assumes the opening content block holds the reply should pick blocks by type instead of by position. Safety classifiers in Haiku 5.5 may also decline requests, and there is no server-side fallback, so clients need to handle stop_reason: "refusal". Organizations that committed to Priority Tier for Haiku 4.5 should note that Haiku 5.5 does not offer it.
On the other hand, Claude Managed Agents users only need to change the model name, according to the guide. In Claude Code, running /claude-api migrate invokes a bundled skill that swaps model IDs and, where needed, handles breaking parameter changes, and it checks the scope with you before editing anything.
Benchmarks, and the line Anthropic draws itself
In the announcement’s table, Haiku 5.5 scores above both Haiku 4.5 and GPT-6 Luna on every evaluation where a value is listed1. A selection:
| Evaluation | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna |
|---|---|---|---|
| GDPval-AA v2.1 | 1620 | 735 | 1437 |
| OSWorld 2.1 (Offline subset) | 72.4% | 15.7% | 48.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% |
| FrontierCode 1.1 (Main) | 46.4% | — | 42.4% |
| Chartography (no tools) | 46.4% | 6.4% | 29.1% |
All of these are figures published by Anthropic, and the table does not say where the GPT-6 Luna numbers come from.
The more useful part is the boundary Anthropic sets for itself. Sonnet 5.5, included in the same table for reference, scores 70.6% on Terminal-Bench 4.0, far above Haiku 5.5’s 39.2%. Anthropic writes that Sonnet 5.5 and Opus 5.5 are still the better fit for complex agentic coding of the kind Terminal-Bench 4.0 measures, while Haiku 5.5 is strongest on tightly scoped jobs such as compaction, summarization, and subagent tasks, which earlier Claude models made too expensive to run. The natural reading is not that Haiku 5.5 replaces the main agent, but that the work delegated underneath a larger model just got cheaper.
The page also carries early-tester feedback. AlphaSense reports a statistically significant gain over Haiku 4.5 across 400 queries, 0.84 versus 0.76, and Box says Haiku 5.5 beat Haiku 4.5 by 11 points while taking about half as long to respond. Both are the companies’ own evaluations.
On safety, the cybersecurity safeguards on Haiku 5.5 are tighter than on Haiku 4.5, though a little more permissive than on other recent Anthropic models. They let through a broader set of defensive work than Sonnet 5.5’s safeguards, yet still stop methods more typical of attackers, penetration testing among them. Its biology safeguards match those of Sonnet 5, Sonnet 5.5, and Opus 5.
Sonnet 5.5 cache reads drop from $0.20 to $0.10
On the same day, Sonnet 5.5’s prompt-cache read price fell from $0.20 to $0.10 per million tokens2. Expressed as a multiple of the base input price, that is a move from 0.1x to 0.05x; cache writes and every other price stay where they were.
Anthropic argues that because cache reads account for a large share of token consumption, the change cuts the cost of most agentic tasks on Sonnet 5.5 by around 20%1. That too is the company’s estimate, and the actual saving likely depends on how much of a given workload hits the cache. When Sonnet 5.5 launched on September 28, its cache-read price was $0.20; input and output rates were left unchanged then, and only the cache side has moved now.
The same announcement also introduced monthly API credits for the Claude Platform for Max and Team subscribers: $100 a month on Max 5x, $200 on Max 20x, and up to $500 on Team, pooled across users.
Availability, and whether to switch now
From launch day, Haiku 5.5 could be used through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS2. It has a 1M-token context window and a 128K-token output limit. According to Anthropic, it is the first Haiku model whose effort level can be tuned. The ID claude-haiku-5-5 is fixed: it carries no date and has no alias.
For pipelines that send large numbers of short prompts, such as summarization or classification, the price gap matters a great deal. But code that uses budget_tokens, non-default sampling parameters, or prefill will start getting 400 errors as soon as the model name is changed. The safer order is to work through the migration guide’s checklist, recount tokens, and only then switch.
Sources
- Introducing Claude Haiku 5.5 - Anthropic official announcement (October 7, 2026)
- Claude Platform release notes - Anthropic official release notes (October 7, 2026 entries)
- Claude Haiku 5.5 migration guide - Anthropic official documentation (checked October 8, 2026)
Was this article helpful?
Thank you!
Received. Thank you!