Overview
Claude Haiku 5.5 is the small model Anthropic released on October 7, 2026, billed as the cheapest and fastest model it has ever shipped. The headline is price: prompts up to 100K tokens cost $0.10 per million input tokens and $0.50 per million output tokens, one tenth of Haiku 4.5 and exactly what OpenAI charges for GPT-6 Luna. Past 100K tokens the rate goes up five times to $0.50 and $2.50, which is still half of what the previous generation charged.
Price is only half of this release. The other half is capability. Haiku 4.5 was close to blank on coding and computer use, scoring 0 on Terminal-Bench 4.0 and 15.7 percent on OSWorld, while Haiku 5.5 brings those to 39.2 percent and 72.4 percent. It is also the first Haiku with an adjustable effort setting, five notches from Low to Max with Medium as the default, so you can spend more compute when the job deserves it. For teams running tens of thousands of calls a day, this is the release that moves small models from acceptable to actually competent.
Key Features
- Pricing split at 100K: Up to 100K tokens it costs $0.10 per million input and $0.50 per million output, with cache reads at $0.01 and cache writes at $0.125, and batch requests at half price. Beyond 100K everything is five times higher. Anthropic says about 90 percent of Haiku 4.5 requests stayed under the line, so most users see the full discount
- Five effort levels: Low, Medium, High, xhigh and Max, defaulting to Medium. Higher settings spend more compute and usually score better while burning more tokens. Artificial Analysis measured two to four extra intelligence points per step at 1.6 to 1.7 times the cost: Medium scores 34 at $0.047 per task, Max scores 43 at $0.213
- 1M context, 128K output: A one million token context window with up to 128K output tokens in standard mode, and 300K in beta for batch. Knowledge cutoff is June 2026. It takes images alongside text
- Coding and computer use catch up: Terminal-Bench 4.0 moves from 0 on Haiku 4.5 to 39.2 percent, and the OSWorld 2.1 offline subset from 15.7 to 72.4 percent. With tools, Humanity's Last Exam goes from 18.7 to 57.4 percent. Anthropic still points to Sonnet 5.5 and Opus 5.5 for complex agentic coding
- Built to work as a subagent: The jobs Anthropic assigns it are high volume, latency sensitive and tolerant of imperfect precision: summarization, compaction, classification, database queries, live support and browser use. It also sits under Opus 5.5 and Sonnet 5.5 as a coding subagent, taking the cheap front-end steps off the expensive model
- Available everywhere at launch: The API model name is claude-haiku-5-5, and it is on Amazon Bedrock, Google Cloud and Microsoft Foundry from day one, plus Claude Code. Every plan from Free to Enterprise can reach it, and Claude Code v2.1.293 makes it the default Haiku model
Use Cases
- Classification, extraction and summarization pipelines that run tens of thousands of times a day and could only be sampled before because of unit cost
- Subagent duty under Sonnet 5.5 or Opus 5.5 so the expensive model only plans and wraps up
- Live customer support and browser use where time to first token matters
- Bulk document compaction and database queries, halved again through the batch API
Pros
- One tenth of the previous generation under 100K tokens, and still half price above the line
- Coding and computer use go from near zero to genuinely usable, which small models could not claim before
- Adjustable effort lets one model cover both the cheap and the careful end of the same workload
- 1M context with 128K output leaves no obvious gap on long documents
- Lands on all three major clouds and Claude Code on launch day, with no channel waiting game
- Subscribers also get monthly API credits, effectively turning subscription fees into spendable budget
Pricing
Billing is per token. Prompts up to 100K tokens cost $0.10 per million input and $0.50 per million output, with cache reads at $0.01 and cache writes at $0.125. Above 100K tokens it is $0.50 and $2.50, with cache reads at $0.05 and cache writes at $0.625. The batch API halves standard rates. Every Claude plan from Free to Enterprise can use it, and Max 5x and Max 20x subscribers receive $100 and $200 of monthly Platform API credits respectively, while Team plans get up to $500 shared across the team. Credits apply to any model on the platform, expire at the end of the month and do not roll over.
Summary
What makes Haiku 5.5 interesting is not how strong it is but how much it can do at this price. Small models used to be a trade where you bought cheap and accepted weak. This time Anthropic matched the competitor on price while fixing the parts that held the line back, so for the first time the small model qualifies for a class of work on its own.
Two things are worth deciding up front. The first is the 100K line. Most traffic stays below it, but if a retrieval flow reliably lands between 120K and 180K tokens the bill will look heavier than the headline suggests, so routing by prompt length pays off. The second is the effort dial. Medium is the default and the value starting point, classification and routing gain little from going higher, while extraction from messy documents and multi-step tool calls are where High and xhigh earn their cost.
The ceiling is real. Factual knowledge is the thin spot, so anything needing encyclopedic accuracy still goes up a tier, and automation benchmarks came in low with evaluators suspecting over-refusal from safety policy that Anthropic says it is fixing. Taken as a primary reasoning model it will disappoint. Taken as the worker model on a pipeline, it is hard to beat on value right now.
Version History
- Haiku 5.5 launch (2026-10-07): Anthropic released Claude Haiku 5.5 as its cheapest and fastest model: $0.10 per million input and $0.50 per million output for prompts up to 100K tokens, one tenth of Haiku 4.5, rising to $0.50 and $2.50 beyond that. It brings a 1M context window, up to 128K output tokens, and the first adjustable effort setting on a Haiku model with five levels from Low to Max defaulting to Medium. Terminal-Bench 4.0 climbs from 0 to 39.2 percent, the OSWorld 2.1 offline subset from 15.7 to 72.4 percent, and the Artificial Analysis Intelligence Index reaches 43 at max effort. Anthropic also halved Sonnet 5.5 cache reads to $0.10 per million and introduced monthly Platform API credits for Max and Team subscribers. Available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Code