AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Oct 1, 2026
·
anthropicclaudeapipricingmigrationbreaking-changesagentstool-useprocurementdevelopers

Claude Sonnet 5.5 costs exactly what Sonnet 5 cost — and breaks five things on the way in, plus a sixth that raises no error

TL;DR: Claude Sonnet 5.5 shipped on 28 September 2026 at $2 / $10 per million tokens — identical to Claude Sonnet 5 on every line of the rate card, including cache and batch rates. The migration instruction Anthropic leads with is a one-string change. Behind it the docs list five breaking changes that turn working code into 400 errors, and a sixth that returns HTTP 200: text the model writes between tool calls now arrives in thinking blocks, and at the default display setting it arrives empty, so a streaming agent UI goes silent mid-task with no error. Effort levels are recalibrated, so carried-over settings no longer mean what they meant. The API default effort is high, while Anthropic’s own agentic guidance says start at medium. A same-price upgrade is not a free upgrade, and the two documents saying so are on the same site.

The cheapest-looking upgrade of the year

There is a category of vendor announcement that platform teams approve without a meeting: the model refresh that is better, faster, and costs the same. Anthropic’s 28 September release of Claude Sonnet 5.5 is textbook. The newsroom line is “a clear upgrade over Sonnet 5 that runs 30% faster and costs up to 30% less for most work.” The pricing page confirms no number moved. The migration guide opens with a code block showing claude-sonnet-5 becoming claude-sonnet-5-5 in seven languages.

Then the migration guide says: “Then check six things.”

That gap — between a one-line diff and a six-item checklist — is the whole story, and it is the same gap this desk keeps finding in different clothes. In September alone: a 40% cost reduction that turned out to be a change in default effort, a refusal policy that started billing for responses containing nothing, an audit log that kept the row and dropped the noun. None of those were undisclosed. All of them were filed under a heading that made them look smaller than they were.

What actually breaks

Anthropic’s what’s-new page is unusually good here: it names the five breaking changes in a bulleted list at the top, before any marketing. Worth reproducing the shape of them, because the list is doing real work.

ChangeFailure modeWho is exposed
thinking: {"type": "disabled"} rejected400 invalid_request_error, message points to between_toolsAnyone who turned thinking off for latency or cost
Forced tool use removed400 on tool_choice any or tool, including on token countingStructured-extraction pipelines, router agents
Thinking blocks bound to model + conversation400 on replay after history edits (accounts created on or after 31 Aug 2026)Anything that rewrites system prompts, trims history, or edits turns
computer_20251124 rejected on Claude API and Google Cloud400 naming the rejected tool typeComputer-use integrations not yet on computer_toolset_20260801
Advisor tool pairings narrowed400 when the advisor is Opus 4.8, Opus 4.7 or Sonnet 5Beta advisor-tool users

Four of those five are loud. A 400 with a descriptive message is the good kind of breaking change: it stops the request, names the problem, and points at the fix. The disabled → between_tools error even tells you what to send instead. Credit where it is due — that is better disclosure hygiene than most of this market manages.

Two details in that table deserve more than a row each.

The computer-use rejection is platform-dependent. On the Claude API and Google Cloud, declaring computer_20251124 returns 'claude-sonnet-5-5' does not support tool types: computer_20251124. On Amazon Bedrock, Claude Sonnet 5.5 accepts the same tool. Identical model, identical request, different outcome depending on which cloud invoices you. Teams running multi-cloud for availability now have a code path that must diverge by provider, which is precisely the kind of thing that gets discovered during a failover rather than before one.

Thinking-block binding is a one-way door. Sonnet 5.5 can read thinking blocks produced by Sonnet 5, Opus 4.8, Haiku 4.5 and earlier models. No model reads Sonnet 5.5’s blocks — not Opus 5.5, not Fable 5.1, not Mythos. So migrating a live conversation onto Sonnet 5.5 preserves its reasoning, and moving off it does not. When a request carries a block the target model cannot read, the API drops it before the model sees it; the request succeeds and the dropped tokens are not billed. The drop is only visible if you send the thinking-binding-controls-2026-08-01 beta header and read the input_transformations array.

Read that as a procurement fact rather than an API detail. Multi-model routing — the standard answer to rate limits, outages and cost control — now silently degrades in one direction. Your fallback path still returns 200. It just returns 200 having thrown away the reasoning that got the conversation that far, and by default it does not tell you.

The change that raises no error

The sixth item is the one to put in front of whoever owns your agent’s user experience.

On Sonnet 5.5, notes the model writes between tool calls come back as progress-update thinking blocks rather than text blocks, once they run longer than a sentence or two. At the default display: "omitted", the text in those blocks is empty. Anthropic’s phrasing is exact and worth quoting: an application that streams those notes to its users “goes quiet between tool calls, with no error.”

Consider what that looks like in production. A user asks an agent to do something that takes ninety seconds and eleven tool calls. On Sonnet 5, they watched a running commentary — checking the schema, now querying the staging table, that returned nothing so trying the replica. On Sonnet 5.5 with no code change, they watch a spinner. The task still completes. The output is still correct. Every metric on your dashboard is green, because nothing failed. What changed is the only thing the user could actually see, and it changed in the direction of looking broken.

The remedy is two lines: set thinking.display to a value that returns the text under adaptive thinking, or switch to between_tools, where the text comes back without any additional setting. The cost is entirely in discovery. A canary that asserts on status codes, latency percentiles and output correctness passes this change perfectly. Catching it requires a person watching a stream — which is exactly the test that gets automated away first.

The arithmetic underneath “30% less”

The saving Anthropic claims is real in kind: a faster model that needs fewer tool calls finishes a task for fewer tokens, and when the per-token rate cannot move, that is the only honest place left to find a discount. But three documented facts sit between the claim and your invoice.

Effort levels are recalibrated. The docs are explicit that an effort level “doesn’t produce the same amount of thinking as it did on Claude Sonnet 5” and instruct teams to “re-run your effort sweep rather than carrying a setting over.” Whatever medium meant to your cost model last week, it means something else now.

The default is not the recommendation. Sonnet 5.5’s default effort on the Claude API is high. Anthropic’s own guidance in the same document: “For agentic coding and multistep tool use, start at medium for well-specified tasks.” A team that changes only the model string lands on high by default and runs agent workloads at a setting the vendor steers away from. That is the Opus 5.5 pattern inverted — there the default was the discount; here the default is the premium.

Independent measurement points the other way. Artificial Analysis puts Sonnet 5.5 at 56 on its Intelligence Index, a strong result, while flagging the model as “very verbose” — roughly 410M output tokens across its evaluation suite against an 81M median for comparable models. Verbosity is not a benchmark footnote when output bills at five times input. It is worth holding that number loosely, since it reflects one harness at one effort configuration and the index has been repriced and restructured recently, but it is the only third-party read available three days in, and it does not corroborate a cheaper bill by default.

The net: “up to 30% less for most work” is a hypothesis about your workload, not a term of your contract. The rate card is the contract, and the rate card did not move.

What did get better, unambiguously

A fair accounting has to include the parts that are straightforwardly good, because there are several and they are the reason to migrate at all.

The knowledge cutoff moves from January 2026 to June 2026 — five months of the world, on a model that is otherwise price-identical. That alone justifies the upgrade for most retrieval-light workloads. The minimum cacheable prompt drops from 1,024 tokens to 512, which brings prompt caching within reach of short system prompts that previously could not use it at all. Per-message effort, mid-conversation system messages and mid-conversation tool changes are all available on Sonnet 5.5 and none of them exist on Sonnet 5 — and the last of those is the direct fix for the append-only constraint that thinking-block binding imposes, since changing instructions with a mid-conversation message rather than an edit keeps the prefix check happy. The tool-use system prompt overhead falls from 354 tokens to 286. And on-demand compaction pairs specifically well with this model, because it is the supported way to shorten a conversation without invalidating the thinking blocks you keep.

That is a genuinely strong release. The complaint here is not about the engineering. It is that the migration cost was priced at one line and documented at six.

The pattern worth naming

Same-price refreshes are becoming the dominant upgrade shape across this market — GPT-6.1 Sol landed at the same $2/$10 a day later, and GPT-6 Sol’s own pricing held flat through a mid-flight fix with no snapshot to pin. The rate card is turning into the least interesting number on a model launch, because vendors have discovered they can ship improvement without touching it and let the token count carry the difference.

For buyers this inverts a long-standing habit. The question at a model refresh used to be what does it cost now. On Sonnet 5.5 that question has a one-word answer — unchanged — and answering it tells you almost nothing. The questions that carry information are what breaks, what changes shape without breaking, how many tokens the same task now consumes, and whether the default configuration is the one the vendor actually recommends. All four are answerable from public documentation, and none of them are on the pricing page.

What to do

Grep for two strings before anything else. "disabled" in your thinking configuration and tool_choice values of any or tool. These are hard 400s, they are trivially findable, and they will take down a production path the moment the model string changes. Replace forced tool use with tool_choice: {"type": "auto"} plus strict tool use, or move the schema into structured outputs.

Audit every place your request builder edits history. System-prompt rewrites, dynamic tool lists, message trimming, custom summarisation — each one now risks a 400 on replayed thinking blocks for accounts created on or after 31 August 2026. Move instruction changes to mid-conversation system messages and keep the transcript append-only. If you need an escape hatch, thinking.block_binding.prefix_mismatch_behavior set to "drop_block" under the thinking-binding-controls-2026-08-01 beta header degrades instead of failing.

Treat model fallback as lossy and instrument it. If your router spills from Sonnet 5.5 to any other model, reasoning is dropped at the boundary with no error and no charge. Send the beta header, log input_transformations, and decide deliberately whether a degraded fallback is better than a queue. Teams standardising on Claude or Claude Code for long agent runs should make that call before the next incident, not during it.

Watch one real agent run with human eyes. Not a unit test, not a status-code canary — a person looking at the stream a customer would see. That is the only reliable way to catch the progress-update change, and developers shipping agent UIs should assume it affects them until they have watched it not happen.

Re-run your effort sweep before you quote a cost. Effort levels mean something new, the default is high, and the vendor’s agentic recommendation is medium. Anyone comparing coding tools or weighing Claude against ChatGPT on price this quarter is comparing rate cards that have quietly stopped being the thing that determines the bill.

Do not rush. Sonnet 5 is marked legacy, not deprecated, with retirement committed no sooner than 30 June 2027. That is nine months of runway to do this deliberately. The one thing not worth doing is what the first line of the migration guide makes look sufficient.

Update, 1 October 2026 — the migration is no longer optional for one set of users. The advice above holds for teams on Sonnet 5, which remains legacy rather than deprecated. It does not hold for Sonnet 4.5. On 30 September Anthropic deprecated claude-sonnet-4-5-20250929 with retirement on 30 November 2026 and named Sonnet 5.5 — this model, with the six-item checklist above — as the recommended replacement. Those teams have 61 days, not nine months, and they face two breaking changes that Sonnet 5 users do not: temperature, top_p and top_k return a 400 from Opus 4.7 onward, and the extended-thinking mode Sonnet 4.5 used is not accepted here at all. The list price also falls from $3/$15 to $2/$10, which looks like a 33% cut until you account for the newer tokenizer producing roughly 30% more tokens for the same text — the effective saving is closer to 13%. The notice-period arithmetic and the three migration targets.

Frequently asked questions

Did Claude Sonnet 5.5 change the price at all?

No. Anthropic's pricing page lists Claude Sonnet 5.5 and Claude Sonnet 5 on adjacent rows with identical figures in every column: $2 per million base input tokens, $2.50 per million for 5-minute cache writes, $4 per million for 1-hour cache writes, $0.20 per million for cache hits and refreshes, and $10 per million output tokens. Batch rates are identical too, at $1 and $5. The what's-new page states it directly — 'Claude Sonnet 5.5 has the same prices as Claude Sonnet 5, including prompt caching and batch processing rates.' The cache multiplier is the standard 0.1x, not the 0.05x that Claude Opus 5.5 gets or the 0.025x on Claude Fable 5.1, so Sonnet 5.5 is the one recent Anthropic release that did not move any number on the rate card. The '30% faster and costs up to 30% less for most work' line in the announcement is a claim about tokens and tool calls consumed per task, not about the rate you are charged per token. Those are different quantities and only one of them is contractual.

What are the five breaking changes, in one place?

First, `thinking: {"type": "disabled"}` returns a 400 `invalid_request_error`; the replacement is `thinking: {"type": "between_tools"}`, which works at `high` effort or below and returns a 400 at `xhigh` or `max`. Second, forced tool use is gone — `tool_choice` set to `{"type": "any"}` or `{"type": "tool", "name": "..."}` returns a 400 reading `tool_choice: type "tool" and "any" are not supported for this model.`, and the same check applies on the token counting endpoint. Third, thinking blocks are bound to the model and the conversation: Sonnet 5.5 reads blocks from Sonnet 5, Opus 4.8, Haiku 4.5 and earlier, but no other model reads Sonnet 5.5's blocks, and on accounts created on or after 31 August 2026 replaying a block after an edit to earlier history returns a 400. Fourth, on the Claude API and Google Cloud the `computer_20251124` computer use tool is rejected in favour of `computer_toolset_20260801` — though Amazon Bedrock still accepts the older tool, so the same code passes on one platform and fails on another. Fifth, the beta advisor tool refuses Opus 4.8, Opus 4.7 and Sonnet 5 as advisors to a Sonnet 5.5 executor.

What is the sixth change, and why does it matter more than the five?

Because it does not fail. On Sonnet 5.5, the notes the model writes between tool calls come back as progress-update `thinking` blocks rather than `text` blocks when they run longer than a sentence or two. At the default `display: "omitted"`, those blocks arrive with their text empty. Anthropic's own wording is that an application streaming those notes to users 'goes quiet between tool calls, with no error.' Every status code is 200. Every field validates. No exception is raised, no alert fires, and no test that asserts on success will catch it. What changes is that the running commentary your users watch during a long agent task — the part that makes a ninety-second tool chain feel alive rather than hung — silently stops arriving. The fix is small once you know: set `thinking.display` to a value that returns the text when using adaptive thinking, or use `between_tools`, where the text comes back without any extra setting. The problem is entirely in the finding, and a staging environment that checks for errors rather than watching the output will not find it.

Is the '30% cheaper per task' claim wrong?

Not wrong, but it is not a discount and it is not guaranteed by anything. It describes fewer tool calls and faster completion on Anthropic's own workload mix, which is a real form of saving and the honest way to improve a model that cannot get cheaper per token. Two things sit awkwardly beside it. The docs state plainly that 'effort levels are recalibrated' and that an effort level 'doesn't produce the same amount of thinking as it did on Claude Sonnet 5,' with the instruction to 're-run your effort sweep rather than carrying a setting over' — so the per-task token count on your workload is an open question until you measure it. And the API default effort is `high`, while Anthropic's own guidance for agentic coding and multistep tool use is to 'start at `medium`.' A team that swaps the model string and changes nothing else is therefore running the configuration the vendor does not recommend for agent work, on a model whose effort levels mean something new. Independent measurement points the same way: Artificial Analysis scores Sonnet 5.5 at 56 on its Intelligence Index but flags it as 'very verbose,' recording roughly 410M output tokens across its evaluation suite against an 81M median. Verbosity is billed at $10 per million.

What should a team running Sonnet 5 in production do this week?

Five things. Grep for `"disabled"` in thinking configuration and for `tool_choice` values of `any` or `tool` — those two produce hard 400s and are the fastest to find and fix, with strict tool use or structured outputs as the replacement for forced calls. Audit whether anything in your request builder edits earlier turns: system-prompt rewrites, tool-list changes, message trimming or history compaction all now risk a 400 on replayed thinking blocks, and the durable fix is to keep conversations append-only and use mid-conversation system messages instead of edits. Check your fallback and routing logic, because thinking-block binding is one-directional — a router that spills from Sonnet 5.5 to Opus 5.5 under load drops the reasoning at the boundary, silently and unbilled. Watch a real agent run end to end with a human looking at the stream, not just at exit codes, to catch the progress-update change. And re-run your effort sweep before you believe any cost figure, including Anthropic's. Sonnet 5 is not going anywhere in the meantime: it is marked legacy, not deprecated, with retirement committed no sooner than 30 June 2027.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.