Anthropic's 80% inference margin went public the same weekend your Claude Code seat got 17% smaller
TL;DR: Three Anthropic stories landed inside 72 hours. On 12 September, Dario Amodei’s essay asked the industry to slow the pace of capability gains. On 13 September, the Financial Times reported — via a shareholder briefing, relayed by Bloomberg — that Anthropic expects a second consecutive quarter of adjusted operating profit, on gross margins above 80% before revenue share and training costs. On 14 September, today, the Claude Code weekly-limit change took effect: a permanent 25% raise over the pre-May baseline, which is a 17% cut against what subscribers had yesterday. Sitting underneath all three is The Information’s 6 September tally of at least 14.8 GW of contracted compute worth up to $517bn, and a listing reportedly targeting ~$2tn. The useful reading is not hypocrisy. It is that the seat limit is the one lever a vendor can pull without publishing a number — and the pre-IPO window is when that lever gets used.
What is new, and what is not
Worth separating, because three of these have been reported at different times and only two are fresh.
The Information’s compute tally is eight days old. Published 6 September, it aggregates announced and contracted deals since October 2025 into at least 14.8 gigawatts and a headline figure of up to $517bn stretching across roughly a decade: about $100bn over ten years with AWS for ~5 GW of Trainium3, roughly $200bn over five years across Google and Broadcom, $50bn with Fluidstack, $45bn with Nscale, $45bn with SpaceX, $35bn with Lambda, around $30bn for ~1 GW of Azure, and the 2 GW AMD MI450 agreement from July. That is context, not news, and it is an outside estimate rather than a company disclosure — an upper bound on contracts that phase in over years, not cash spent or a liability due now. Prior investor guidance was about $180bn on server rental through 2029.
What is genuinely new arrived on 13 and 14 September:
- Anthropic told a small group of shareholders it expects adjusted operating profit this quarter, a second consecutive profitable period on that measure, with gross margins above 80% before revenue-sharing payments to partners such as Amazon and before model training costs (FT, via Bloomberg, 13 September).
- Reporting from 12 September onward puts the IPO at a target valuation around $2tn — against May’s $965bn Series H post-money — a raise of up to roughly $100bn, a pricing window before the November midterms, and Nvidia in talks for an anchor stake of up to $10bn. None of it company-confirmed.
- The Claude Code limit change took effect today, 14 September.
The margin and the limit cut are the two that a buyer has to act on, and they point the same direction.
May established the pattern. September broke it.
On 6 May, Anthropic did something unusually legible: it announced a SpaceX Colossus 1 compute deal and doubled Claude Code rate limits in the same week, removing peak-hour throttling for Pro and Max at the same time. Capacity in, limits up. The causal story was told out loud, and it was easy to believe because the arithmetic was intuitive: GPUs are scarce, more GPUs means more headroom, headroom gets passed to subscribers.
Run the same arithmetic on today. Contracted capacity is now an order of magnitude larger than it was in May. Inference margins are reportedly north of 80%. The company expects its second straight adjusted-profit quarter. And the subscription allocation went down 17%.
Both of those can be true because the May story was never really about physics. It was about a promotion that a well-capitalised private company could afford to run. The 50% temporary boost in force since May is now expiring, and what replaces it — 125 against a pre-May base of 100 — is genuinely higher than the baseline and genuinely lower than what heavy users have been working with all summer. Anthropic’s own reposted clarification put it plainly: “compared to today, this works out to a 17% reduction in weekly limits on Claude Code.”
The part that has not changed since we first wrote it up yesterday is the part that should govern planning. Anthropic has never published the absolute weekly token figure that 100 refers to — not before the promotion, not during it, not now. Percentages of an unpublished base are not a capacity commitment. They are a description of a capacity commitment, issued by the only party who can see it.
The margin explains the lever
Here is why the FT number makes the limit change more predictable rather than more outrageous.
If inference gross margin is above 80%, then the marginal cost of serving a subscription seat’s tokens is small relative to the seat price — but the variance is not. A heavy Claude Code user on a flat monthly fee is the classic buffet problem: the price is fixed, the consumption is not, and the tail is expensive. The three ways to manage that are to raise the price, to meter the usage, or to cap it. Raising the price is visible and comparable against Cursor, Codex and Copilot. Metering means publishing a rate card, which invites the same comparison. Capping a percentage of an undisclosed base is the only one of the three that does not create a number anyone outside can benchmark.
That is not a conspiracy; it is just the cheapest available lever. And the reason to expect it to be pulled again is the calendar. A company preparing to list — Anthropic filed its confidential S-1 on 1 June, ten days after OpenAI — is a company whose margin stops being a private matter and becomes a quarterly commitment. The $10.9bn revenue and $559m adjusted operating profit projected for Q2, which we covered in May, worked out to roughly a 5.1% operating margin. That is a thin cushion. Anything that widens it without touching a published price will be attractive right through the listing window.
It is also not the first unilateral revision this quarter. In August, Priority Tier — the only Anthropic tier carrying a published uptime target — was marked no longer available for purchase, after a month with 21 logged incidents. In September, the Fable 5.1 cache-read price cut changed which workload shapes are cheap without changing the headline rate. The pattern across all three is the same: the published number stays still and the thing behind it moves.
What to do about it
The action is not “leave Anthropic.” Claude Code remains, on capability, one of the strongest coding harnesses available, and the Claude models behind it are not what changed this week. The action is to stop treating a subscription seat as a capacity plan.
Split the workload by whether you need to forecast it. Interactive, supervised, bursty work — a developer at a terminal — belongs on a seat, where hitting a wall costs an afternoon. Scheduled agents, CI jobs, large batch refactors and anything with a delivery date belong on metered API access, where the rate card is published per token and a vendor change shows up as a cost variance instead of a hard stop. That split is cheap to implement and it is the single change that makes the next limit revision a non-event.
Instrument what you actually consume. Most teams that were surprised today were surprised because they did not know how close to the ceiling they were running. Log per-developer weekly consumption now, while the new limit is fresh, so the next change can be evaluated against real numbers rather than forum anecdotes.
Keep a second harness warm, not theoretical. A configured, credentialled alternative running a small share of real work — Cursor, Codex, Copilot, whichever fits the stack — converts a future terms change from a migration project into a config flag. The cost is a few hours a quarter. The thing it hedges is not this 17%; it is the class of change that produced it.
Ask the portability question in writing. As noted in today’s earlier piece on the standards-body reporting, a vendor commitment you can read, version and diff is worth more than one you can only infer. For Claude Code specifically, the question worth putting in a renewal thread is simple: what absolute weekly token figure does our plan entitle us to, and how much notice do we get before it changes? A vendor that will not answer the first half has told you what to plan for.
The bottom line
None of the three stories this weekend is scandalous on its own. A CEO can argue for slower capability growth while buying inference capacity. A company can run a promotion and end it. A margin above 80% on the marginal token is a good business, not a betrayal.
What they add up to is a change in how the vendor relationship should be modelled. Until this weekend, Anthropic’s subscription limits could plausibly be read as a function of available capacity, because in May the company said so itself. After this weekend — 14.8 GW contracted, 80%-plus inference margins, a second profitable quarter, a $2tn listing in the window — that reading no longer holds. Subscription capacity at Anthropic is a pricing instrument, priced against a base only Anthropic can see, tuned on a schedule set by a prospectus. Plan the work you have to deliver on a meter you can read.
Frequently asked questions
Is the 80% gross margin figure Anthropic's real margin?
No, and the qualifier is the whole point. The Financial Times reporting, summarised by Bloomberg on 13 September 2026, is that Anthropic's gross margins exceed 80% before revenue-sharing payments to channel partners such as Amazon and before the cost of training models. Those are two of the largest cost lines in frontier AI. What the figure describes is the unit economics of serving a token once the model exists and once the distribution deal is set aside — roughly, the margin on inference. That is a genuinely useful number, because it is the one that governs whether a vendor has room to cut prices or reason to tighten capacity. It is not a company-wide profit margin, and anyone quoting it as one is misreading it. Anthropic has not published the figure itself; it reportedly came from a shareholder briefing, and the company has not publicly confirmed the financial details.
Does the 17% Claude Code cut mean Anthropic is short on compute?
Nothing in the public record supports that reading, and a fair amount cuts against it. The Information's 6 September tally put Anthropic at a minimum of 14.8 gigawatts of contracted capacity since October 2025, with a headline value of up to $517 billion over roughly the next decade — against prior investor guidance of about $180 billion in server rental through 2029. Meanwhile inference gross margins are reportedly above 80% and the company expects a second consecutive quarter of adjusted operating profit. A vendor that is capacity-starved and margin-thin behaves differently from one that is contracting gigawatts and clearing 80% on the marginal token. The more consistent explanation is that the subscription tiers were running a promotional allocation since May and the promotion ended. That is a pricing decision, which is a perfectly ordinary thing for a company to make — it just should not be planned around as though it were physics.
Why does the IPO timing change how I should read any of this?
Because it changes what the number is for. Anthropic confidentially filed a draft S-1 on 1 June 2026. Reporting since 12 September points to a listing targeting a valuation in the region of $2 trillion — against the $965 billion post-money of May's Series H — a raise reported at up to roughly $100 billion, pricing sought before the November US midterms, and Nvidia in talks for an anchor position of up to $10 billion. None of that is confirmed by the company. But a margin that is about to be printed in a prospectus and then re-tested every quarter by public-market analysts is a margin under permanent pressure, in a way that a private company's margin is not. Practically: assume the subscription tiers will keep being tuned through the listing window, and do not build a 2027 capacity plan on a 2026 promotional allocation.
What should a team actually change this week?
Move the workloads you must be able to forecast onto a meter whose rate card is public. Anthropic publishes per-token API prices; it has never published the absolute weekly token figure that Claude Code's Pro, Max, Team and seat-based Enterprise limits are a percentage of. That asymmetry is the practical finding, not the direction of this particular change: you cannot forecast a percentage of a number you cannot see, and the vendor can move it with a blog post. Keep subscription seats for interactive, bursty, individually-supervised work where hitting a wall is an inconvenience. Put CI jobs, batch refactors, scheduled agents and anything with a deadline on metered API access, or on a second harness, so a limit change is a cost variance rather than an outage.
Is a second vendor worth the switching cost just to hedge a rate limit?
Hedging a rate limit alone probably does not justify it. Hedging the whole class of unilateral terms changes usually does. In the past four months the same vendor has withdrawn the only tier carrying a published uptime target, run a large temporary limit increase and then retired it, and revised a limit announcement after developers challenged the framing. Each event on its own is minor. Together they describe a supplier whose service definition moves faster than a procurement cycle, which is normal for this market and not specific to Anthropic. The cheap version of the hedge is not a full migration: keep one alternative harness configured and credentialled against the same repository, run a small share of real work through it continuously so it is known to work, and make sure your prompts and tooling are not written against one vendor's harness conventions. That converts a future terms change from a project into a config flag.
Did Amodei's pacing essay and the spending numbers actually contradict each other?
Less than the headlines suggest, and the distinction matters for procurement. The essay published on 12 September argues for slowing the rate at which model capabilities improve, coordinated across labs and backed by evaluation. Contracting gigawatts is about serving existing models to a growing user base — inference, not frontier capability. A company can coherently believe the industry should slow capability growth while buying capacity to serve demand it already has. Where the tension is real is in the incentives: a listed company with a $2 trillion valuation to defend and a decade of contracted compute to fill has a structural reason to keep capability demand growing, whatever its stated position. That is an argument for reading vendor safety commitments as documents with version numbers rather than as dispositions, which is the same conclusion the standards-body reporting pointed to.
Sources
- Bloomberg — Anthropic Sees Adjusted Operating Profit This Quarter, FT Says (13 September 2026)
- Stocktwits — Anthropic IPO: Claude Maker Reportedly Targets Second Straight Quarter Of Adjusted Profit (13 September 2026)
- Investing.com — Anthropic wants AI to slow down. Its $517 billion spending plan says otherwise (14 September 2026)
- Data Center Dynamics — Anthropic signed $517bn in compute agreements in past 11 months (The Information tally, 6 September 2026)
- 24/7 Wall St. — Anthropic Locks Down $517 Billion in Compute Ahead of IPO (13 September 2026)
- Benzinga — Nvidia Eyes Up To $10 Billion Investment in Anthropic IPO at $2 Trillion Valuation (12 September 2026)
- BleepingComputer — Anthropic is cutting Claude Code's current weekly limits by 17%
- Dario Amodei — We Must Pace the Frontier (12 September 2026)
- Anthropic — Pricing (API rate card)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.