AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Sep 28, 2026
·
anthropicclaudepricingapisafetyclassifiersprocurementplatform-riskbillingdevelopers

Anthropic now bills refusals that return nothing — in the three categories where its own docs say benign work gets caught

TL;DR: On 24 September 2026 Anthropic resumed billing for Claude refusals that arrive before any output, in three of five categories: bio, frontier_llm and reasoning_extraction. cyber, general_harms and null-category refusals remain free. A billed refusal charges input tokens at the full rate of the model that ran it, returns an empty content array, and still counts against your rate limits. With fallback enabled, you now pay for both the refused attempt and the retry. The stated rationale — disrupting coordinated probing at scale — is sound. The friction is in the selection: Anthropic’s own category table says “beneficial life sciences work can also trigger” bio and “benign machine learning work can also trigger” frontier_llm, while reasoning_extraction describes a prompt shape, not a harm. The reassurance figure — 99.7% of Claude Code, Claude.ai and Cowork accounts saw no billable block — measures subscription surfaces, not the metered API organisations the change actually bills. No per-category false-positive rate has been published, so the exposure cannot be priced in advance.

The change, in one sentence

Anthropic’s platform release notes for 24 September 2026 record it plainly: the company is “resuming billing for refusals that arrive before any output when stop_details.category is "bio", "frontier_llm", or "reasoning_extraction", the categories where we measure low volumes of false positives.”

Until that date, a refusal that produced nothing was free. Claude’s safety classifiers had already read your input tokens, and Anthropic absorbed that cost. Now, in three categories, it does not. The refusal docs spell out the mechanics: these refusals “are billed like any other request, at the rates of the model that ran it,” and “either way, content is empty and token counts appear in usage. The request still counts against your rate limits.”

Start with the case for the change, because it is genuinely strong. Anthropic says it acted after observing coordinated attacks on its systems. A refusal that costs nothing is a probe that costs nothing, and an adversary trying to map the boundary of a classifier — or to extract capability through thousands of slightly varied attempts — has every incentive to keep firing. Metering that loop is a rational defence, and it is the same logic that prices any other abusable free resource. Anthropic also disclosed the change in its release notes and wrote it into the reference documentation with a dedicated table column rather than burying it. That is better behaviour than the industry baseline.

The question worth asking is not whether to bill probes. It is who else is standing in the blast radius.

One note for anyone triaging that release note as a whole: the billing change was not the only breaking item in it. The same 24 September entry also records that the Compliance API Activity Feed stopped returning filenames, project document names and artifact titles — retroactively, on events recorded months earlier. A team that reads the release note for pricing exposure and stops there will miss an audit-log change that affects a different function entirely.

The three categories, read against Anthropic’s own descriptions

Here is the category table as the documentation presents it, with the billing column intact:

categoryWhat Anthropic says it meansBilled before any output
cyberCould enable cyber harm. “Benign cybersecurity work can also trigger this category.”No
bioCould enable biological harm. “Beneficial life sciences work can also trigger this category.”Yes
frontier_llmCould assist development of competing AI models. “Benign machine learning work can also trigger this category.”Yes
reasoning_extraction”The request asks the model to reproduce its internal reasoning in the response text.”Yes
general_harmsUsage-policy areas outside the four named categories. “Benign work can also trigger this category.”No

Read the middle column and the right column together. For two of the three billed categories, Anthropic’s own one-line description contains an explicit warning that legitimate work sets them off. The company is not hiding this; it is documented in the same table as the price. But it means the new charge lands on a population that includes, by the vendor’s own account, computational biologists, drug-discovery teams, ML researchers, evaluation engineers and anyone fine-tuning or benchmarking models — the people whose prompts routinely read like the thing the classifier is looking for.

The third billed category is the odd one. reasoning_extraction is not a harm description at all. It describes a request that asks the model to write out its internal reasoning, and points you at adaptive thinking as the structured alternative. That is a prompt-hygiene note wearing a safety classifier’s uniform. Instructions of the form show your work and explain your reasoning step by step have been standard prompt-engineering advice for years and are embedded in system prompts, eval harnesses and agent scaffolds written long before preserved thinking existed. Those prompts now cost money to have declined.

There is a further asymmetry there. Anthropic’s default fallback routing carries documented recommendations for cyber, bio and frontier_llm, and the docs note that “for categories with no recommended fallback, the refusal stands.” So a reasoning_extraction refusal is billed, empty, and has nowhere to fall through to.

Fallback used to soften this. Now it compounds it.

This desk’s coverage of the Opus 5.5 launch set out the fallback economics as they stood on 22 September: a flagged request refuses visibly on the API, and if you opt into fallbacks: "default", the retry runs on a model Anthropic recommends for that category — Opus 4.8 for cyber, Opus 5 for bio and frontier_llm, both at $5/$25 against Opus 5.5’s $4/$20. The counterintuitive finding then was that the fallback costs more per token than the model you chose.

The 24 September change adds a second charge underneath the first. The refusal documentation now states: “When you use fallback, the refusal that triggered it is billed, in addition to the fallback request, when it arrived mid-stream or is in one of the billed categories.”

So a life-sciences team running Opus 5.5 with fallback enabled, on a request the bio classifier declines, now pays:

  1. The refused Opus 5.5 attempt — input tokens at $4/MTok, zero output, empty response.
  2. The Opus 5 fallback attempt — input at $5/MTok and output at $25/MTok.

Before 24 September, only line 2 existed. The usage.iterations array is the per-attempt record, and the docs are clear that “tokens from different models are never summed into one field” — which is precisely why a team watching only top-level usage counts will not see line 1 appear. That is the same class of measurement gap this desk flagged when Anthropic’s usage fields stopped describing the bill: the number your code reads and the number on the invoice have drifted apart again, in a new place.

One genuine mitigation ships alongside. Sticky routing records which model served a conversation after a fallback and sends later turns in that conversation directly to it, scoped to your organisation and retained for roughly an hour. The documentation’s stated purpose is to avoid “paying for an attempt that would predictably be declined again on every turn.” On a long flagged conversation, that is the difference between paying the refusal toll once and paying it on every turn.

But sticky routing arrives with server-side fallback, and server-side fallback is beta on the Claude API only — the docs note it “is not available on Amazon Bedrock, Google Cloud, or Microsoft Foundry.” The billing change, by contrast, “applies on all platforms.” Teams on Bedrock, Vertex or Foundry get the charge without the server-side mitigation and must configure the SDK middleware for client-side fallback instead. That asymmetry is worth a line in a platform-selection memo.

What the 99.7% figure measures

Anthropic’s public reassurance is that 99.7% of Claude Code, Claude.ai and Cowork accounts encountered none of these newly billable blocks during testing. Taken at face value, that is a low number and a fair thing to publish.

It also measures the wrong population. Claude.ai and Cowork are subscription products; a user there is not billed per token and cannot receive a refusal charge in the sense this change creates. Much Claude Code usage likewise runs on a subscription plan rather than metered API billing. The change applies to metered requests — on the Claude API and on all three partner clouds. So the denominator is dominated by accounts that the mechanism does not reach, and it is silent on the accounts it does: API organisations, whose traffic skews toward exactly the specialised technical workloads the three billed categories exist to catch.

The figure a buyer would need is the per-category false-positive rate on API traffic, and it has not been published. Nor has a refund path for a block that turns out to be wrong. The reporting on the change has converged on this same gap: without a refusal frequency, a false-positive rate, or an estimate of monthly impact, a developer cannot price the change before adopting it.

This is the recurring shape of 2026 pricing. The rate card is not where the cost lives. It lives in a default, a classifier threshold, or a billing rule in a docs page — the same pattern as effort defaults doing the work a discount was credited for, a cache-read cut that only pays out on the right workload shape, and a classifier that is free until a gateway drops a field.

The structural read

Two years of frontier-model safety policy have converged on a pattern this desk has tracked repeatedly: the rate card buys you a restricted model, and full capability is a clearance rather than a line item. Life sciences work runs through a verification programme with its own retention terms; capability tiers across the frontier labs are now gated by access programmes rather than price; independent evaluation itself has become a funded relationship rather than an open one.

What changed on 24 September is that the unverified state acquired a price. Previously, being on the wrong side of a classifier cost you capability and round trips. Now, in three categories, it costs capability, round trips and money — and the meter runs whether the classification was right or wrong. That asymmetry is the whole of the objection. An abuser and a misclassified researcher pay identical amounts, and only one of them has a way to stop paying.

None of which makes the change wrong. Metering a probe loop is defensible, and 0.3% is a real number even if it describes the wrong denominator. The reasonable position is narrower: a charge that lands on false positives should be accompanied by a published false-positive rate and a refund path, and at the moment it is accompanied by neither.

What to do

Measure your actual exposure before reacting to it. Query your logs for stop_reason: "refusal" across the last 90 days, group by stop_details.category, and price the input tokens in the three billed categories at your model’s input rate. Most teams will find a number in the low single-digit dollars. Knowing that is worth more than assuming it, and it is the same discipline — cost per completed task, not cost per million tokens — that every other 2026 pricing change has demanded.

Fix reasoning_extraction in your prompts, not your budget. Grep system prompts, agent scaffolds and eval harnesses for instructions asking the model to narrate its reasoning in the response text, and move to adaptive thinking. This is the one billed category where an edit removes the exposure outright, and it is invisible until someone goes looking. Teams running Claude Code or custom harnesses on top of the Claude API should treat this as a one-hour audit.

Branch on stop_reason, and read usage.iterations. A refusal is an HTTP 200 with an empty content array. Code that branches only on content length sees a silent null result, and code that reads only top-level usage will not see the refused attempt it just paid for. Both are now billing bugs as well as correctness bugs — a point that applies to any developer shipping against a classifier-gated model, and one worth weighing when comparing Claude against ChatGPT for a workload that sits near a policy boundary, since OpenAI’s cyber-tier models gate capability by programme too.

Treat the billed-category column as a variable. The documentation says outright that the set “may change as Anthropic keeps measuring and refining its safeguards’ false positive rates.” cyber and general_harms are free today because the measured false-positive volume is higher, not because they are exempt in principle. Anyone budgeting security or general-purpose work on the assumption that refusals are free is budgeting against a configuration setting, not a commitment.

Frequently asked questions

Which refusals are billed now, and which are still free?

Anthropic's refusal category table carries a column headed 'Billed before any output', and it splits the five categories three-to-two. Billed: bio (the request could enable biological harm), frontier_llm (the request could assist development of competing AI models, restricted under Anthropic's commercial terms) and reasoning_extraction (the request asks the model to reproduce its internal reasoning in the response text). Not billed: cyber and general_harms, along with any refusal whose category comes back null — which the documentation describes as a normal, permanent value rather than a placeholder. Two details are easy to miss. First, mid-stream refusals were always billed and still are: if the model produced any output before declining, you pay for the input tokens and the output already streamed, in every category. This change is specifically about refusals that arrive before a single output token. Second, the split is explicitly provisional. The docs state that 'the billed categories may change as Anthropic keeps measuring and refining its safeguards' false positive rates', which means the column is a live configuration setting rather than a contract term, and the cyber row could move to 'Yes' without a new pricing page.

What does a billed refusal actually cost?

Input tokens at the full rate of the model that ran the request, with output tokens at zero because nothing was generated. The arithmetic is unremarkable on a short prompt and unpleasant on a long one. On Claude Fable 5.1 at $10 per million input tokens, a 50,000-token research context that trips the bio classifier costs $0.50 for an empty content array. On Claude Opus 5.5 at $4 per million, the same context costs $0.20. Prompt caching softens this considerably — a cache hit on Fable 5.1 is $0.25 per million tokens and on Opus 5.5 it is $0.20 per million — so a refused turn deep inside a warm agent loop is cheap, while a refused first turn on a cold, uncached 200,000-token document set is not. Two non-price costs travel with the charge. The response still counts against your rate limits, so a run of refusals burns quota as well as budget. And when server-side fallback is enabled, the documentation is explicit that 'the refusal that triggered it is billed, in addition to the fallback request' whenever the refusal is in a billed category — so a flagged bio request now bills twice, once for the empty Fable or Opus attempt and once for the fallback model that answers.

Why is reasoning_extraction the strangest of the three?

Because it is not a description of harm, it is a description of a prompt shape. Anthropic defines it as a request that 'asks the model to reproduce its internal reasoning in the response text', and points developers at adaptive thinking to get reasoning in a structured form instead. Nothing in that definition involves misuse. It covers a category of prompt that has been standard practice since chain-of-thought prompting entered general use: instructions along the lines of show your work, walk through your reasoning step by step, or explain how you arrived at that. Those strings sit in production system prompts, in evaluation harnesses, in agent scaffolds and in prompt libraries across the industry, most of them written before preserved thinking existed as an alternative. A team that inherits such a prompt now pays the full input rate every time it fires and gets an empty response. It is also the category with the weakest recovery path: Anthropic's default fallback routing has documented recommendations for cyber, bio and frontier_llm, and the docs note that 'for categories with no recommended fallback, the refusal stands'. So the practical shape of a reasoning_extraction refusal is billed, empty, and not automatically retried anywhere. The fix is a prompt edit rather than a procurement decision, but it is a prompt edit nobody has been told to make.

Does the 99.7% figure mean my exposure is negligible?

It means less than it appears, because of which population it measures. Anthropic's reassurance, as reported at the time of the change, is that 99.7% of Claude Code, Claude.ai and Cowork accounts encountered none of these newly billable blocks in testing. Read the three surfaces named. Claude.ai and Cowork are subscription products where a user does not see a per-token invoice at all, and a large share of Claude Code usage runs on a subscription plan rather than metered API billing. The billing change applies to metered requests, on every platform including Amazon Bedrock, Google Cloud and Microsoft Foundry. The figure therefore describes a population that substantially overlaps with the people the change cannot bill, and it is silent on the population it can: API organisations, which skew heavily toward exactly the specialised workloads — computational biology, model evaluation, fine-tuning research, automated red-teaming — that the three billed categories are built to catch. This is not an accusation of bad faith; it is a mismatch between a metric and a question. The number that would settle the question is the false-positive rate per category on API traffic, and no such figure has been published. As one write-up put it, there is no figure for how often refusals happen, no false positive rate, and no estimate of what this adds to a monthly bill, so a developer cannot price the change in advance.

What should a team actually do this week?

Five things, in rough order of payoff. First, query your logs for responses with stop_reason set to 'refusal' over the last 90 days, group by stop_details.category, and multiply the input-token counts in the three billed categories by your model's input rate — that is your retroactive exposure estimate, and for most teams it will be reassuringly small, which is worth knowing rather than assuming. Second, grep your system prompts and agent scaffolds for reasoning-disclosure instructions and replace them with adaptive thinking; this is the one category where a text edit removes the exposure entirely. Third, if you are on the Claude API, enable server-side fallback for its sticky routing, which records the model that served a conversation and sends later turns straight there — the docs describe this as avoiding 'paying for an attempt that would predictably be declined again on every turn', and on a repeatedly flagged conversation it is the difference between paying the refusal toll once and paying it every turn. Fourth, note the platform asymmetry: server-side fallback is not available on Bedrock, Google Cloud or Microsoft Foundry, so teams on those platforms get the new charge without the server-side mitigation and must use the SDK middleware to get client-side fallback instead. Fifth, if your legitimate work trips bio or frontier_llm regularly, the structural answer is a verification programme rather than a budget line, and the application is worth starting before the invoice arrives.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.