AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Oct 8, 2026
·
anthropicclaudeapipricinghaikulong-contexttokenizercost-modellingdevelopersagents

Claude Haiku 5.5 is a tenth of Haiku 4.5's price — until your prompt passes 100,000 tokens, where it costs five times more

Anthropic’s release notes for 7 October 2026 carry two lines. The first launches Claude Haiku 5.5 — claude-haiku-5-5, a 1M-token context window, 128K maximum output, adaptive thinking on by default — at $0.10 input and $0.50 output per million tokens. That is a tenth of the rate on Haiku 4.5, the model it supersedes, and it makes Haiku 5.5 the cheapest model Anthropic has ever shipped by a wide margin.

The second line lowers the prompt cache read price on Claude Sonnet 5.5 from $0.20 to $0.10 per million tokens, a move from the standard 0.1x multiplier to 0.05x of base input. Cache writes and every other rate are unchanged.

Neither line is the story. The story is a footnote on the pricing page: Haiku 5.5 appears there twice.

Anthropic’s first prompt-length price tier

ModelInputOutputCache readBatch in/out
Haiku 5.5, prompt ≤100,000 tokens$0.10$0.50$0.01$0.05 / $0.25
Haiku 5.5, prompt >100,000 tokens$0.50$2.50$0.05$0.25 / $1.25
Haiku 4.5 (previous generation)$1$5$0.10$0.50 / $2.50
Sonnet 5.5 (for comparison)$2$10$0.10$1 / $5

Per million tokens, from Anthropic’s pricing documentation, verified 8 October 2026.

Every rate multiplies by exactly five across the threshold — input, output, cache read, cache write, batch. And the multiplier applies to the whole prompt, not to the tokens above the line. A prompt of 100,000 tokens costs $0.01 in input. A prompt of 100,001 tokens costs $0.05. One extra token quintuples the input bill.

Put the other way round: for what Anthropic charges to process a 100,001-token prompt, the cheap tier would have processed 500,005 tokens. Crossing the line costs you the price of 400,000 tokens you did not send.

The steepest long-prompt surcharge in the market, at the lowest threshold

Length-tiered pricing is not new. What is new is the shape of Anthropic’s version.

VendorThresholdInput aboveOutput above
Anthropic, Haiku 5.5100,0005x5x
Google, Gemini 3.1 Pro and 2.5 Pro200,0002x1.5x
xAI, Grok 4.7 and 4.3200,0002x2x
OpenAI, GPT-6 and GPT-5.6 families272,0002x1.5x

Anthropic’s threshold is half Google’s and xAI’s, and 37% of OpenAI’s. Its multiplier is two and a half times theirs on input and more than three times theirs on output. On the two axes that matter — where the cliff is and how far you fall — Haiku 5.5 is the harshest long-prompt rate card any major vendor currently publishes.

It is also, within Anthropic’s own lineup, unique. The pricing page is explicit that Claude 4.6 and later models except Haiku 5.5 include the full 1M-token context window at standard pricing: a 900,000-token request on Sonnet 5.5 bills at the same per-token rate as a 9,000-token one. Opus 5.5, Sonnet 5.5 and Fable 5.1 have no length tier at all. Anthropic has built a model that is sold on a 1M-token window and priced so that 90% of that window costs five times the advertised rate.

The tokenizer moves the cliff 23% closer

Here is the part that will surprise teams migrating from Haiku 4.5, and it is the reason to measure before you route traffic.

Anthropic’s pricing page states that Claude 4.7 and later models use a newer tokenizer producing approximately 30% more tokens for the same text, and that Sonnet 4.6 and earlier models use the previous one. Haiku 4.5 is on the old tokenizer. Haiku 5.5 is on the new one.

The 100,000-token threshold is denominated in Haiku 5.5’s tokens. So in the units your existing budget is written in:

100,000 ÷ 1.30 ≈ 76,900 Haiku-4.5-equivalent tokens.

A prompt that metered at 85,000 tokens on Haiku 4.5 — comfortably inside any 100,000-token mental model — meters at roughly 110,500 on Haiku 5.5 and lands on the expensive side of the line. Teams that sized their context windows against the old tokenizer and assumed the threshold gave them headroom will find they are already over it on day one.

What the migration actually saves

Take a representative cached agent turn: 2,000 fresh input tokens, 500 output tokens, and a resident cached context. Price it on both models, applying the 1.3x tokenizer inflation to the Haiku 5.5 counts.

Short-prompt workload — 10,000 old tokens of context, 13,000 new:

Long-prompt workload — 150,000 old tokens of context, 195,000 new, so the high tier applies:

So the honest summary of the launch is: Haiku 5.5 cuts short-prompt costs by about 87% and long-prompt costs by about 35%. Both are real savings. Only one of them resembles the headline.

Above the cliff, the cheap model stops being the cheap option

The Sonnet 5.5 cache cut on the same day is what makes this interesting rather than merely annoying.

Below the threshold, Haiku 5.5 reads cached tokens at $0.01 against Sonnet 5.5’s new $0.10 — a tenfold gap. Above the threshold, Haiku 5.5 reads at $0.05 against the same $0.10 — a twofold gap. For a cache-dominated agent, which is what most production agents are, the entire economic case for choosing the small model loses 80% of its force the moment the context crosses 100,000 tokens.

The same turn as above, this time against Sonnet 5.5, with 95,000 tokens of cached context (under the line) and then 150,000 (over it):

Resident cached contextHaiku 5.5Sonnet 5.5Haiku advantage
95,000 tokens$0.0014$0.018513.2x cheaper
150,000 tokens$0.00975$0.02402.5x cheaper

Two readings, both actionable. Letting an agent’s context drift from 95,000 to 150,000 tokens makes the Haiku turn seven times more expensive for 58% more context. And at 150,000 tokens you are no longer choosing between a cheap model and an expensive one — you are paying half of Sonnet 5.5’s price for a materially weaker model, which is a trade many teams would decline if the rate card framed it that way.

The corollary is the most valuable line in the whole launch. Trimming a 120,000-token cached context to 95,000 takes the cache-read cost of a turn from $0.006 to $0.00095 — an 84% cut for 21% less context. There is no other configuration change available on Anthropic’s platform with that ratio.

The migration notes, briefly

Haiku 5.5 is not a drop-in replacement for Haiku 4.5. Anthropic documents three breaking changes: manual extended thinking via budget_tokens returns a 400, adaptive thinking is on by default so responses can begin with thinking blocks, and the same text counts as more tokens. The model is live on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry, and is listed Active with a retirement floor of not sooner than 7 October 2027.

Haiku 4.5 remains Active, at $1/$5, with its own floor of 15 October 2026 — a floor, not a date, and no deprecation notice has been issued. Anthropic’s 60-day notice commitment makes a retirement this month impossible in principle.

What to do

  1. Measure, in the right units. Count your resident context with the token-counting endpoint against claude-haiku-5-5, not against an old model and not from a characters-÷-4 estimate. The answer you want is whether you are above or below 100,000 of its tokens.
  2. If you are between roughly 77,000 and 200,000 old tokens, trim. That band is where the cliff was invisible under the old tokenizer and where the payoff for cutting context is largest — up to 84% of a cached turn.
  3. If you cannot trim, re-price against Sonnet 5.5, not Haiku 4.5. Above the line the gap is 2x on cached input, which is usually not enough to justify the capability difference.
  4. Fix your cost dashboard’s assumptions. Any model that bills by prompt length breaks per-token cost averages, because the average silently depends on a distribution you are not tracking. Bucket your requests either side of 100,000 tokens.

Our API cost calculator applies each vendor’s long-prompt multiplier to the whole request when the input crosses its threshold, and each vendor’s published characters-per-token figure on top, which is what makes these comparisons come out differently from a straight rate-card division. The full lineup and plan arithmetic is in the Claude review; dates and floors for every model are in the retirement calendar, and the dated price history is in the pricing tracker.

Frequently asked questions

Is Claude Haiku 5.5 cheaper than Haiku 4.5 or not?

Cheaper in both tiers, but by wildly different margins. Below 100,000 tokens the rate is $0.10/$0.50 against Haiku 4.5's $1/$5 — a tenfold cut on the sticker, realising as about 7.7x once you account for Haiku 5.5 running the newer tokenizer that produces roughly 30% more tokens for the same text. Above 100,000 tokens the rate is $0.50/$2.50, a twofold sticker cut that realises as about 1.5x after the tokenizer. So the honest range is: expect to pay about an eighth of your Haiku 4.5 bill on short-prompt traffic and about two thirds of it on long-prompt traffic. Nobody loses money migrating; the question is whether you get the saving you budgeted for.

What exactly counts toward the 100,000-token threshold?

Prompt length — the input side of the request, which on a conversational or agentic workload means the entire resident context: system prompt, tool definitions, conversation history, documents and cached prefixes. It is not the output and it is not a per-turn increment. The practical consequence is that a long-running agent session crosses the line once and then stays across it, with every subsequent turn repriced at five times the rate, including the cache reads that make up most of such a turn. There is no API flag and no opt-out: the tier is a function of what you sent.

Does the 5x surcharge apply to Sonnet 5.5, Opus 5.5 or Fable 5.1?

No. This is the most useful thing to know about it. Anthropic's pricing page states that Claude 4.6 and later models, with Haiku 5.5 as the stated exception, include the full 1M-token context window at standard pricing — a 900,000-token request bills at the same per-token rate as a 9,000-token one. Haiku 5.5 is the only model in Anthropic's lineup with a prompt-length price tier. So above 100,000 tokens the comparison that matters is no longer Haiku 5.5 against Haiku 4.5; it is Haiku 5.5 against Sonnet 5.5, where the price gap has narrowed to roughly 2x on cached input.

Why does the cliff arrive sooner than my token budget predicts?

Because the threshold is denominated in Haiku 5.5's tokens, and Haiku 5.5 uses Anthropic's newer tokenizer. Anthropic's own pricing documentation says that tokenizer produces approximately 30% more tokens for the same text, and that Sonnet 4.6 and earlier — which includes Haiku 4.5 — use the previous one. Dividing through, 100,000 new tokens correspond to roughly 76,900 old ones. A prompt that metered at 85,000 tokens on Haiku 4.5 meters at about 110,500 on Haiku 5.5 and lands on the expensive side of a line it never approached before. Anthropic notes the exact increase depends on content and workload shape, so measure your own corpus rather than trusting the 1.3 multiplier as a constant.

What changed about Sonnet 5.5's pricing on the same day?

Anthropic lowered the prompt cache read price on Sonnet 5.5 from $0.20 to $0.10 per million tokens, moving it from the standard 0.1x multiplier to 0.05x of base input. Cache writes and every other rate are unchanged, so Sonnet 5.5 remains $2/$10. Two consequences. First, Sonnet 5.5's launch pitch — identical to Sonnet 5 on every line of the rate card — has lapsed, in the buyer's favour: Sonnet 5 still reads cache at $0.20. Second, Sonnet 5.5 now matches GPT-6.1 Sol's $0.10 cached-input rate, removing the one line on which the OpenAI model undercut it, though Anthropic's tokenizer still counts more tokens for the same text.

What should I actually do this week?

Three things, in order. Measure the current resident context of anything you plan to route to Haiku 5.5, in Haiku 5.5 tokens via the token-counting endpoint rather than in old token counts or character estimates. If it sits between roughly 77,000 and 200,000 old-tokenizer tokens, trimming it below the new 100,000-token line is the single highest-leverage cost change available — our worked example cuts a cached turn by 84% for 21% less context. And if trimming is not possible because the workload genuinely needs the context, price it against Sonnet 5.5 rather than against Haiku 4.5, because above the cliff you are choosing between a weak model at half the price and a strong one at double, not between cheap and expensive.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.