Insights

News & commentary

Claude Haiku 5.5: sticker price down 90%, the real bill needs a recount

On October 7, 2026, Anthropic released Claude Haiku 5.5, billed as its cheapest, fastest and most capable small model yet, aimed at high-volume, cost-sensitive work. The headline is price: $0.10 input and $0.50 output per million tokens for prompts under 100K tokens, down 90% from the previous generation. The fine print matters more: a new tokenizer counts roughly 30% more tokens for the same text, five legacy integration patterns now return HTTP 400, and adaptive thinking is on by default. Meanwhile the industry's price war has moved to the cheap tier, with DeepSeek, Google and Anthropic each pricing it differently.

SlateMoth Editorial · Muse ·

Muse research and drafting; Muse staged review in the same author context. Researched, drafted, fact-reviewed and translated into eight languages by Muse in staged same-author review (reviewMode=muse_same_platform_staged, contextRelationship=same_author_context); disclosed as same-author context, not an independent third-party audit. Prices and benchmarks are vendor-reported; migration details come from official docs; third-party comparisons are attributed with limitations.

What happened

On October 7, Anthropic released Claude Haiku 5.5, describing it as its cheapest, fastest and most capable small model to date. The positioning is explicit: high-volume, cost-sensitive work - summaries, context compaction, database queries, classification, and subagent duty under Opus 5.5 and Sonnet 5.5.

Price is the headline: for prompts up to 100K tokens, $0.10 per million input tokens and $0.50 per million output tokens; above 100K tokens the rates rise to $0.50 and $2.50. The previous Haiku 4.5 charged flat rates of $1.00 and $5.00. Anthropic says around 90% of requests to the old Haiku fell into the cheap tier.

[1]

Reading the price cut: a two-tier sticker

By Anthropic's account, Haiku 5.5 costs around 75% less to run on average than 4.5. Note the attribution: this is a vendor claim built on its own traffic assumptions. Your actual bill depends on how long your prompts are, how much caching you use, and the token inflation from the new tokenizer.

Two more value changes shipped alongside: cache reads on Sonnet 5.5 were halved in price, which Anthropic says makes most agentic work about 20% cheaper; and Max and Team subscribers get new monthly API credits - $100 for Max 5x, $200 for Max 20x, up to $500 pooled for Team - meant for experimenting with agents on the platform.

Haiku 5.5 is also the first Haiku with an adjustable effort setting: users can lean toward saving money or toward more intelligence within a single request instead of switching model tiers. It is the first time a small model gets the thinking throttle previously reserved for flagships.

[1]

Benchmarks: how to read a vendor scorecard

The scorecard is Anthropic's own reporting: on knowledge work, GDPval-AA v2.1 scores 1620 (735 for 4.5, 1437 for GPT-6 Luna); AA-Briefcase v1.1 scores 1578 (614 / 1336); on computer use, OSWorld 2.1 offline subset reaches 72.4% (15.7% / 48.9%); agentic coding via Terminal-Bench 4.0 reaches 39.2% (0.0% for 4.5, 16.4% for Luna); visual reasoning via Chartography reaches 46.4% (6.4% / 29.1%).

How to read it: the biggest jumps are in hands-on tasks - OSWorld from 15.7% to 72.4%, Terminal-Bench from zero to 39.2%. But keep two caveats: every number is vendor-reported and not independently reproduced; and Anthropic itself says complex agentic coding remains Sonnet and Opus territory, with Haiku suited to narrow, repetitive work.

[1]

Migration: five places your old code returns 400

Beyond price, the real cost hides in the migration guide. Haiku 5.5 switches to the newer tokenizer shared with Claude 4.7 and later models: the same text counts roughly 30% more tokens than on 4.5. Every token-denominated budget, truncation threshold and cost dashboard has to be recomputed, using counts from the new model, not recycled old numbers.

Harder still are five outright 400 errors: temperature must be 1, top_p must be 0.99, any top_k is rejected; assistant prefill is refused; manual thinking budgets (budget_tokens) are rejected; computer use must move to the new toolset; and sending thinking blocks back after editing earlier turns also returns 400. The old classifier recipe of temperature=0 plus prefill plus max_tokens=20 fails at every step.

Then the silent changes: adaptive thinking is on by default, so responses may open with a thinking block that carries an empty field and only a signature; thinking tokens count toward max_tokens, so small limits can truncate right after thinking; and a new safety classifier can stop a request with a refusal stop reason with no server-side fallback. Migration is not a model ID swap; it is a small refactor.

[2] [4]

The cheap tier becomes a three-way race

Zoom out: between September and early October 2026, the cheap tier shipped in a five-week burst - DeepSeek V4.1 Flash, Google Gemini 3.8 Flash and Anthropic Haiku 5.5 all landed. The flagship headlines gave way to the budget tier: for chatbots, ticket routers and extraction jobs called tens of thousands of times a day, the bill is the real benchmark.

The three vendors picked three pricing logics: DeepSeek charges by peak and off-peak hours, doubling at peak; Anthropic charges by prompt length, quintupling past 100K tokens; Google charges a flat rate. One third-party comparison puts Haiku 5.5's entry-tier output price at roughly one seventh of Gemini 3.8 Flash - but there is no universal cross-vendor pricing formula, and changing caching, timing or length assumptions changes the conclusion.

One price alignment is worth noting: several outlets report Haiku 5.5's short-prompt pricing matches OpenAI's GPT-6 Luna exactly. But Anthropic's own announcement never published Luna's prices, so the parity claim comes from third-party reporting and should not be cited as official confirmation.

[3] [2] [4]

Conclusion and action list

The conclusion: model competition in 2026 has entered the total-cost-of-ownership phase. Tokenizer inflation, default thinking overhead, cache hit rates and prompt length distribution jointly determine a bill; looking at dollars per million tokens alone misleads budgets. Anthropic itself positions Haiku as the fast layer in an agent system - routing, classification, extraction, subagents - leaving complex coding to Sonnet and Opus.

An action list for teams: first recount real prompt distributions with count_tokens under the new tokenizer; set effort explicitly to low for classifiers and replace prefill plus temperature=0 habits with structured outputs or enum tools; run 4.5 and 5.5 side by side on small traffic and compare label agreement before full cutover; and for long-prompt work, check whether you cross the 100K-token line where prices quintuple.

[1] [2]

Sources and further reading

Source records are supplied and reviewed by Muse in the same author context; they have not been independently fact-checked.

  1. Anthropic official announcement (2026-10-07)

    Anthropic

    Recorded publication date ·

    Recorded verification time ·

  2. Anthropic Claude Platform migration docs

    Anthropic (Claude Platform Docs)

    Source publication date not verified

    Recorded verification time ·

  3. Tech Insider cheap-tier comparison (2026-10-08, third party)

    Tech Insider

    Recorded publication date ·

    Recorded verification time ·

  4. DEV Community engineer migration notes (2026-10-08, third party)

    DEV Community

    Recorded publication date ·

    Recorded verification time ·