Skip to content

Claude Haiku 5.5 Launches at Up to 90% Below Haiku 4.5 Prices With a 1M Token Window, but Anthropic Says It Is Not a Frontier Model

Anthropic's Claude Haiku 5.5 launched October 7 at $0.10/$0.50 per million tokens with a 1M context window. Prompts over 100K tokens cost five times more, and its own system card says it is not a frontier model.

Anthropic's official Claude Haiku 5.5 announcement graphic

Anthropic released Claude Haiku 5.5 on October 7, calling it “the cheapest, fastest, and most capable small model we’ve ever released.” It went live on the API, Amazon Bedrock, Google Cloud and Microsoft Foundry the same day under the model ID claude-haiku-5-5.

The headline numbers are real. List price starts at $0.10 per million input tokens and $0.50 per million output tokens, a 90% cut from Haiku 4.5 for most requests, and the context window grows from 200,000 to 1 million tokens.

Some of the hype around the launch goes further than Anthropic does, though. Its own system card says Haiku 5.5 “is not at the capability frontier,” and the 1M window comes with a price jump once a prompt passes 100,000 tokens. Here is what Anthropic actually confirmed.

What Anthropic announced

The @claudeai announcement on X went up at 11:31 PM IST on October 7. It says Haiku 5.5 costs “around 75% less to run than Claude Haiku 4.5” on average. The post had passed 4 million views by midday on October 8.

Anthropic pitches it for high-volume, cost-sensitive work: summaries, compaction, database queries, classification, live customer support and browser use. It also says Haiku 5.5 “pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work,” meaning the bigger model plans and Haiku does the quick lookups.

It is the first Haiku model with an adjustable effort setting. Adaptive thinking is on by default, and developers pick low, medium, high, xhigh or max effort to trade cost against quality. The default is medium.

The price, and the 100,000 token catch

Anthropic’s model overview lists two price bands. Prompts up to 100,000 tokens pay $0.10 input and $0.50 output per million tokens. Prompts over 100,000 tokens pay $0.50 input and $2.50 output, five times more.

Anthropic says the cheaper band covers “around 90% of requests to our previous Haiku model.” So the 1M context window is real, but filling it is not cheap by Haiku standards. Cache reads cost $0.01 per million tokens in the lower band and $0.05 above it, and Batch API requests get 50% off input and output.

There is a second catch in the docs. Haiku 5.5 uses the newer tokenizer that arrived with Claude 4.7, so “the same text counts as approximately 30% more tokens than on Claude Haiku 4.5.” Anthropic says its 75% average saving already accounts for that, but any budget you built in tokens needs recounting.

How it compares with Haiku 4.5, Sonnet 5.5 and Opus 5.5

All figures below come from Anthropic’s launch post and model documentation.

Model Input / output per 1M tokens Context window Max output Anthropic’s latency label
Claude Haiku 5.5 $0.10 / $0.50 (up to 100K prompt), $0.50 / $2.50 (over 100K) 1M 128K Fastest
Claude Haiku 4.5 $1.00 / $5.00 200K 64K Not in current comparison
Claude Sonnet 5.5 $2 / $10 1M 128K Fast
Claude Opus 5.5 $4 / $20 1M 128K Moderate

On benchmarks, Anthropic compared Haiku 5.5 with Haiku 4.5, OpenAI’s GPT-6 Luna and Sonnet 5.5. These are Anthropic’s own reported runs, not independent tests.

Benchmark Haiku 5.5 Haiku 4.5 GPT-6 Luna Sonnet 5.5
GDPval-AA v2.1 (knowledge work, Elo) 1620 735 1437 1840
AA-Briefcase v1.1 (knowledge work) 1578 614 1336 1824
OSWorld 2.1, offline subset (computer use) 72.4% 15.7% 48.9% 83.9%
Humanity’s Last Exam, no tools 45.9% 10.2% Not reported 56.9%
Humanity’s Last Exam, with tools 57.4% 18.7% Not reported 64.5%
Terminal-Bench 4.0 (agentic coding) 39.2% 0.0% 16.4% 70.6%
FrontierCode 1.1 Main (agentic coding) 46.4% Not reported 42.4% 52.1% (xhigh)
Chartography, no tools (visual reasoning) 46.4% 6.4% 29.1% 61.6%

The jump over Haiku 4.5 is large on every row. But Sonnet 5.5 still leads everywhere, and on Terminal-Bench 4.0 the gap is 31 points. Anthropic says so itself: “Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks.”

The GPT-6 Luna column is the more interesting one. Luna has the same $0.10 / $0.50 headline list price, as we covered in our look at GPT-6 Sol and Luna API pricing, and Haiku 5.5 beats it on every benchmark both companies’ models appear in. That is a vendor chart, so wait for independent results before switching.

Is it “ultra-fast”? What the speed claim rests on

Anthropic calls Haiku 5.5 “our fastest model to date.” A footnote narrows it: that holds “at each model’s standard speed,” and Haiku 5.5 “runs less quickly than our Opus models in Fast Mode.”

Anthropic did not publish a tokens per second figure or a time to first token number. The docs only rank it “Fastest” on a comparative latency scale. The concrete speed numbers come from customers. Asana says it saw “over a 30% reduction in latency” and “up to 2.5x faster inference per agent turn” compared with the model it uses today, which it did not name. Box says Haiku 5.5 scored 11 points higher than Haiku 4.5 “at about half the latency.”

Not a frontier model, by Anthropic’s own account

Scores like those make it tempting to call Haiku 5.5 a frontier model. Anthropic does not. The Haiku 5.5 system card says it “is not at the capability frontier” and is less capable than Claude Opus 5 on most evaluations. It even describes the document as one of its condensed cards for “non-frontier models.”

Under its Responsible Scaling Policy, Anthropic says Haiku 5.5 does not cross the CB-2 or Autonomy-2 thresholds, but it treats the model as meeting CB-1 and Autonomy-1 and applies those mitigations.

The card also lists weak spots. Haiku 5.5 “over-refused more than any other model we tested” in Anthropic’s automated behavioral audit. It used a leaked answer without telling the user 17% of the time, up from 2% for Haiku 4.5. And it “hallucinated more than other recent models, and about as much as Claude Haiku 4.5.” On the plus side, Anthropic calls it its most robust Haiku-class model against prompt injection.

Confirmed vs not confirmed

Confirmed by Anthropic Not confirmed or not stated
Released October 7, 2026; model ID claude-haiku-5-5 (Bedrock: anthropic.claude-haiku-5-5) Which Claude app subscription plans can pick it
1M token context window and 128K max output (300K on the Batch API in beta) Any tokens per second or time to first token figure
$0.10 / $0.50 up to 100K tokens; $0.50 / $2.50 above Independent benchmark results
Reliable knowledge cutoff of June 2026 That it is a frontier model; the system card says it is not
Available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS India-specific pricing; API billing is in USD

Anthropic has published the system prompt Haiku 5.5 uses on claude.ai and the Claude apps, but the launch post does not say which plans get it. GitHub separately added Haiku 5.5 to Copilot for Pro, Pro+, Max, Business and Enterprise users, with a gradual rollout.

What breaks when you migrate

Swapping the model ID is not enough. Anthropic’s migration guide lists several breaking changes from Haiku 4.5:

  • Setting temperature, top_p or top_k to a non-default value returns a 400 error.
  • Assistant message prefill returns an error. Requests must end with a user turn.
  • The manual budget_tokens thinking mode is gone. Use adaptive thinking and effort instead.
  • Responses can begin with thinking blocks, and thinking tokens count toward max_tokens, so a small limit can cut off the answer.
  • Safety classifiers can decline a request with stop_reason: "refusal", and there is no fallback model.
  • Priority Tier is not supported on Haiku 5.5.

Teams moving up the range should also read our breakdown of Claude Sonnet 5.5, whose cache reads Anthropic halved to $0.10 per million tokens in the same announcement.

What it means for Indian developers

For startups paying for API calls in dollars, this is a meaningful cost cut. Take a support bot handling 1 million requests a month at about 2,000 input tokens and 500 output tokens each. On Haiku 4.5 list prices that is about $4,500 a month. On Haiku 5.5, after adding roughly 30% for the new tokenizer, it comes to about $585, or roughly ₹56,600 at about ₹96.8 to the dollar (approximate, before GST and card fees).

The savings disappear if your prompts are long. Retrieval apps that stuff 150,000 tokens of documents into each call land in the higher band, where Haiku 5.5 is only half the price of Haiku 4.5, not a tenth. Trimming context below 100,000 tokens is now worth real money.

Plan for the refusal behaviour too. Customer support in Indian languages, fintech KYC flows and health queries are areas where an over-cautious model frustrates users. Log every stop_reason: "refusal" during testing before you route production traffic to it. And if you are choosing between the two cheapest big-lab models, run Haiku 5.5 and GPT-6 Luna on your own evals. They cost the same on paper.

For heavier work, our look at Claude Opus 5.5 covers when the top model is worth its $4 / $20 price.

FAQ

What is the Claude Haiku 5.5 model ID?

On the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS it is claude-haiku-5-5. On Amazon Bedrock it is anthropic.claude-haiku-5-5. Anthropic says it is a fixed ID with no date suffix and no separate alias.

Is the 1M token context window available to everyone?

Anthropic lists 1M tokens as the standard context window for Haiku 5.5 with no beta label. Prompts over 100,000 tokens are billed at $0.50 input and $2.50 output per million tokens, five times the lower band.

How much does Claude Haiku 5.5 cost?

$0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Batch API requests get 50% off. Anthropic says it costs about 75% less to run than Haiku 4.5 on average.

Is Claude Haiku 5.5 better than Sonnet 5.5?

No. Sonnet 5.5 scores higher on every benchmark in Anthropic’s launch table. Haiku 5.5 is the faster, cheaper option for narrow, high-volume tasks.

What is Claude Haiku 5.5’s knowledge cutoff?

Anthropic lists a reliable knowledge cutoff and training data cutoff of June 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *

Loading the next article…

Continue reading