Verified October 7, 2026
Anthropic released Claude Haiku 5.5 on October 7, 2026. The new small Claude model is aimed at high-volume, latency-sensitive workloads including classification, extraction, routing, summarization and subagent tasks.
The headline specifications are a one-million-token context window, up to 128,000 standard output tokens, adaptive thinking, and API pricing starting at $0.10 per million input tokens and $0.50 per million output tokens. The Claude API model ID is claude-haiku-5-5.
Haiku 5.5 is materially cheaper than Haiku 4.5 while gaining a much larger context window and adaptive reasoning. The lowest $0.10/$0.50 prices only apply when the prompt is no larger than 100,000 tokens.
Claude Haiku 5.5 specifications
| Specification | Claude Haiku 5.5 |
|---|---|
| Release date | October 7, 2026 |
| Claude API ID | claude-haiku-5-5 |
| Context window | 1,000,000 tokens |
| Standard max output | 128,000 tokens |
| Batch max output | up to 300,000 tokens, beta |
| Thinking | Adaptive |
| Default effort | medium |
| Input, prompts up to 100K | $0.10 / MTok |
| Output, prompts up to 100K | $0.50 / MTok |
| Input, prompts over 100K | $0.50 / MTok |
| Output, prompts over 100K | $2.50 / MTok |
| Cache read, up to 100K | $0.01 / MTok |
| Batch API | 50% discount on input and output |
| Knowledge cutoff | June 2026 |
The biggest change is the price
Claude Haiku 4.5 was listed at $1 per million input tokens and $5 per million output tokens. Haiku 5.5 cuts those list prices to $0.10 input and $0.50 output when prompts stay at or below 100,000 tokens.
On a pure price-per-token basis, that is a 90% reduction.
Example: 5 million input tokens + 1 million output tokens
For a workload consuming five million input tokens and one million output tokens, with every request remaining below the 100K prompt threshold:
- Haiku 4.5: 5 × $1 + 1 × $5 = $10
- Haiku 5.5: 5 × $0.10 + 1 × $0.50 = $1
That is a nominal 90% difference.
Anthropic says typical tasks cost roughly 75% less to run than Haiku 4.5. That lower real-world figure matters because Haiku 5.5 uses a newer tokenizer. Anthropic’s documentation says the same text can count as roughly 30% more tokens than it did on Haiku 4.5.
For production systems, cost per successful task is therefore a better metric than sticker price per million tokens.
Two pricing tiers
The lowest rates only apply if the prompt contains 100,000 tokens or fewer.
Above that threshold, pricing rises to:
- $0.50 / MTok input
- $2.50 / MTok output
- $0.05 / MTok cache read
This distinction is important because Haiku 5.5 supports a one-million-token context window. A 500K-token prompt is not billed at the same rate as a 20K-token prompt.
The context window grows from 200K to 1M
Haiku 5.5 increases the context window from 200,000 to 1,000,000 tokens, a fivefold jump. Standard maximum output also doubles from 64K to 128K tokens.
That makes the model particularly relevant for:
- large document sets,
- retrieval and RAG systems,
- code repositories,
- long agent histories,
- context compression,
- high-volume document processing.
Anthropic also documents a beta Batch API configuration that can support up to 300K output tokens.
Adaptive thinking reaches the Haiku tier
Haiku 5.5 supports Anthropic’s current adaptive thinking system, with medium as the default effort level.
Instead of assigning every request a fixed reasoning budget, the model can spend different amounts of reasoning depending on the task. That is potentially valuable in agent pipelines: simple routing calls can remain cheap, while a more difficult tool-use or coding subtask can receive more reasoning without immediately escalating to a larger model.
This also affects migration. Integrations built around manual extended thinking or budget_tokens should be updated rather than assuming Haiku 5.5 behaves like Haiku 4.5.
Anthropic positions Haiku 5.5 as a high-volume worker
Anthropic explicitly describes Haiku 5.5 as a model for high-volume, latency-sensitive tasks such as:
- classification,
- extraction,
- routing,
- summarization,
- database queries,
- live support,
- in-app assistants,
- voice agents,
- context compression,
- subagent work.
That positioning matters more than whether Haiku beats Sonnet on every benchmark. A low-cost worker model does not need to solve every open-ended problem. It needs to solve a very large number of narrow and medium-complexity tasks reliably, then escalate difficult cases to a stronger model.
Benchmarks: a large launch-day jump, with an evidence caveat
Anthropic reports large improvements over Haiku 4.5. Two of the most striking published results are:
| Benchmark | Haiku 5.5 | Haiku 4.5 |
|---|---|---|
| OSWorld 2.1, reported offline subset | 72.4% | 15.7% |
| Terminal-Bench 4.0 | 39.2% | 0.0% |
Those numbers point to substantial improvements in computer use and agentic coding.
They should still be framed correctly. The benchmark projects themselves are externally developed, but the Haiku 5.5 scores shown at launch are primarily being communicated by Anthropic. Broad independent reruns using matched settings are naturally limited on release day.
What these scores do not tell you
A synthetic benchmark does not directly measure:
- number of tool calls per successful task,
- retry loops,
- stability over long agent runs,
- real tokens-per-second throughput,
- cache effectiveness in your own prompts,
- fallback frequency to a larger model,
- final cost per successfully completed task.
For an agent platform, cost per successful task is usually the more useful business metric.
Prompt caching can be extremely cheap
For prompts in the lower pricing tier, a cache read costs only $0.01 per million tokens.
That can matter substantially for repeated context such as:
- large system prompts,
- project rules,
- static documentation,
- reusable repository context,
- stable agent background instructions.
A system with a high cache-hit ratio may therefore achieve an effective input cost well below the normal uncached input rate.
Migrating from Haiku 4.5 requires more than changing the model ID
Existing integrations should be tested.
Anthropic highlights changes including:
- Switch the model ID to
claude-haiku-5-5. - Move from manual extended thinking to adaptive thinking.
- Remove unsupported non-default sampling parameters.
- Do not rely on assistant prefill.
- Parse response blocks defensively.
- Review updated computer-use toolsets.
- Re-measure token counts and costs because of the newer tokenizer.
That last point matters for observability. A cost dashboard calibrated to Haiku 4.5 token ratios can systematically misestimate Haiku 5.5 usage.
Availability
Anthropic lists Haiku 5.5 across:
- Claude Platform
- Amazon Bedrock
- Google Cloud
- Microsoft Foundry
- Claude Platform on AWS
Platform-specific model identifiers can differ. In the Claude API, the model ID is claude-haiku-5-5.
Because the model launched on October 7, 2026, cloud-provider documentation and SDK registries may temporarily lag behind Anthropic’s own model documentation.
Claude Haiku 5.5 vs Haiku 4.5
| Category | Haiku 4.5 | Haiku 5.5 |
|---|---|---|
| Context | 200K | 1M |
| Max output | 64K | 128K |
| Thinking | Extended | Adaptive |
| Small-prompt input | $1/MTok | $0.10/MTok |
| Small-prompt output | $5/MTok | $0.50/MTok |
| Cache read | $0.10/MTok | from $0.01/MTok |
| Positioning | fast small Claude | high-volume routing and subagents |
| Tokenizer | older generation | newer; about 30% more tokens for identical text according to Anthropic |
Haiku 5.5 is therefore more than a routine version bump. Anthropic changes the pricing model, context size, reasoning system, tokenizer and intended agent role at the same time.
Who should care about Haiku 5.5?
Strong fit
The model looks particularly attractive when applications generate large numbers of model calls:
- AI agents
- multi-agent systems
- coding subagents
- support automation
- email and document classification
- structured extraction
- summarization
- RAG
- routing
- browser and computer-use subtasks
Less obvious fit
For long, difficult and open-ended coding or planning tasks, a stronger Sonnet or Opus model may remain the better choice.
Haiku 5.5’s real advantage is not “maximum intelligence”. It is the combination of low price, low latency and a sufficiently high success rate.
Bottom line: Lower prices and 1M context for agent workloads
Claude Haiku 5.5 is a major update to Anthropic’s small-model tier: a 1M context window, adaptive thinking, much stronger published launch benchmarks and API pricing starting at $0.10/$0.50 per million tokens.
For short prompts, the nominal price per token is 90% below Haiku 4.5. Real task-level savings are less extreme because of the new tokenizer and workload differences; Anthropic puts typical savings at roughly 75%.
That makes Haiku 5.5 particularly interesting for agent and API-heavy systems. The key question is whether those low token prices translate into a lower cost per successful task in real workloads.
What AI on Mac should test next
A useful first-party benchmark would run the same agent tasks on Haiku 5.5, Haiku 4.5 and a stronger Claude model while measuring:
- success rate,
- total tokens,
- cache-read tokens,
- tool calls,
- latency,
- retries,
- escalations to stronger models,
- cost per successful task.
That would provide a reason to revisit the article beyond the launch-day specification sheet.
Transparency
Sources and review basis
These primary and reference sources form the basis of the technical assessment. Vendor claims and external benchmarks are identified as such in the article.