AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Aug 22, 2026
·
deepseekopenaianthropicpricingchinaagents

The AI price war just flipped: DeepSeek is raising prices up to 1,100% while OpenAI and Anthropic cut

TL;DR: The AI price war just inverted. On 13–14 August 2026 DeepSeek shipped the official build of its flagship V4-Pro and, the same day, said it is raising API prices by 50% to 1,100% with peak/off-peak surge pricing from 17 August. At peak, V4-Pro runs about 9 yuan (≈$1.3) input and 27 yuan (≈$3.8) output per million tokens — roughly a 4× jump on output — while the “1,100%” figure is the cache-hit input rate specifically. In the same window, US labs cut: OpenAI trimmed GPT-5.6 Luna 80% to $0.20/$1.20, Anthropic launched Claude Opus 5 at $5/$25 — half of Fable 5’s $10/$50 — and Google put Gemini 3.7 Flash out at $0.75/$3.75, half of 3.6 Flash. Financial Times data shows US model prices have fallen materially since mid-July, narrowing the US-vs-China gap to roughly 19% on input, 14% on output. The takeaway for buyers: the reflex that “DeepSeek = cheapest” is now wrong for its flagship — GPT-5.6 Luna and Gemini Flash undercut V4-Pro, and only DeepSeek’s V4-Flash ($0.14/$0.28) is still rock-bottom. The durable move is a model-agnostic router, because the “cheapest strong model” title now changes hands monthly.

What just happened

For most of the last two years, AI budgeting ran on one assumption: American frontier models were the expensive, best-in-class option, and Chinese open-weights models — DeepSeek above all — were the cheap fallback you reached for when the bill got scary. In the space of about a week in mid-August 2026, that assumption broke in both directions at once.

On 13–14 August, DeepSeek released the production build of its flagship, DeepSeek-V4-Pro, and simultaneously announced a sweeping change to its API pricing. Per Caixin Global, which first reported the change, rates are rising by 50% to as much as 1,100% depending on the model, the token type and the time of day, and the new prices take effect 17 August 2026. The most-quoted “1,100%” number is the increase on cache-hit input tokens — the cheapest line item — which jumps from roughly 0.025 yuan to about 0.3 yuan per million. The headline workloads move less dramatically but still sharply: at peak, uncached input runs about 9 yuan (≈$1.3) and output about 27 yuan (≈$3.8) per million tokens, roughly quadrupling V4-Pro’s previous $0.435/$0.87. DeepSeek is also introducing peak/off-peak surge pricing — peak windows of 09:00–12:00 and 14:00–18:00 Beijing time, with off-peak roughly half price.

At the same time, the American labs kept moving the other way:

Financial Times data cited across the coverage shows the prices customers actually pay for leading US models have declined materially since mid-July, and the gap between US closed-source pricing and the Chinese market has narrowed to roughly 19% on input and 14% on output. Put the two trends together and you get the story: the cheap option is getting more expensive, and the expensive option is getting cheaper.

The new pricing map

Here is where the major API models sit as of mid-August 2026, per each vendor’s published pricing and the reporting above. All figures are per million tokens, input / output, in USD; DeepSeek’s new rates are converted from yuan and are approximate.

ModelTierInputOutputNote
Anthropic Fable 5Frontier$10$50Most expensive frontier flagship
OpenAI GPT-5.6 SolFrontier$4$20Cut 21 Aug; promo rate expires 21 Nov 2026
Anthropic Claude Opus 5Frontier$5$25Launched 24 Jul at half of Fable 5
Moonshot Kimi K3Large open-weight~$15Cheapest “large frontier-class” option
OpenAI GPT-5.6 TerraMid$2$12Cut ~20% in late July
Anthropic Claude Sonnet 5Mid$2$10Intro rate; Sept rise reportedly scrapped
DeepSeek V4-Pro (from 17 Aug, peak)Flagship≈$1.3≈$3.8Was $0.44/$0.87; off-peak ≈ half
Google Gemini 3.7 FlashWorkhorse$0.75$3.75Intro through 31 Dec, then $1.50/$7.50
OpenAI GPT-5.6 LunaCheap$0.20$1.20Cut 80% in late July
DeepSeek V4-FlashCheap$0.22$0.66Off-peak; doubles at peak. Corrected 22 Aug

Two things jump out. First, DeepSeek’s flagship is no longer a budget product: at peak, V4-Pro’s ≈$3.8 output lands level with Gemini 3.7 Flash’s $3.75 and more than triples GPT-5.6 Luna’s $1.20 — the two American tiers buyers actually reach for on price. The only DeepSeek line that still sits near the price floor is V4-Flash, though DeepSeek’s live pricing page now puts it at $0.22/$0.66 off-peak and $0.44/$1.32 at peak — the $0.14/$0.28 figure we published on 14 August is no longer what the vendor charges. Second, the American mid-and-cheap tiers — Luna, Terra, Sonnet 5, Gemini Flash — now occupy the value band that Chinese models owned a year ago.

Why the flip is happening

The two directions have different causes, and both are rational.

The US labs are buying share with price because capability leads have stopped lasting. When a frontier lead evaporates within weeks — Grok 4.6 matching GPT-5.6 on the intelligence index at launch, Z.ai’s GLM-5.3 and DeepSeek resetting the open-weights floor monthly — “we have the smartest model” is no longer a durable pitch. So the labs compete on the number buyers can actually control: cost per unit of useful work. Anthropic’s Opus 5 is the cleanest example — it did not lower the sticker of its top tier, it shipped more capability at the Opus price band while halving the flagship rate. OpenAI’s Luna cut and Google’s Flash discount are the same move aimed at the highest-volume workload in AI: coding agents that burn tokens by the million.

DeepSeek is raising prices because its flagship finally costs money to run — and because near-zero margin is not a business. The official V4-Pro is a stronger, more expensive-to-serve model than the preview it replaces. DeepSeek’s answer is not a flat hike but surge pricing: charge more when GPUs are saturated (Chinese business hours) and less when they are idle. That is a capacity-management tool as much as a revenue one, and it signals that even the poster child of “too cheap to meter” AI has hit the same infrastructure wall everyone else has. The cheap tier moving far less than the flagship is the tell: DeepSeek is segmenting, pushing price-sensitive traffic to Flash and monetising performance on Pro.

Why this matters

1. The “DeepSeek is cheapest” reflex is now a budgeting bug. If you route bulk traffic to a DeepSeek flagship endpoint on the assumption that it is automatically the low-cost option, you are, from 17 August, very possibly overpaying versus GPT-5.6 Luna or Gemini 3.7 Flash — especially during European daytime, which sits squarely in DeepSeek’s peak window. Re-run the math per workload; the map above is the starting point, your own token mix is the answer.

2. Surge pricing punishes geography. Time-of-day pricing is new to this market, and it is not neutral. DeepSeek’s peak hours (09:00–12:00 and 14:00–18:00 Beijing) overlap European mornings and afternoons almost perfectly, while US business hours fall in the cheap off-peak band. If you are an EU team defaulting to DeepSeek for cost, you are now paying peak rates during exactly the hours you work. This is the same policy we flagged as pending on V4-Flash in early August — it is now confirmed and dated.

3. Cheap is no longer a moat, so lock-in is the real risk. The uncomfortable truth of this market is that the value crown rotates monthly: Luna in late July, Gemini Flash in mid-August, whatever ships next month. Optimising your architecture around this week’s cheapest strong model is how you get stranded when it re-prices — and note that Gemini Flash’s $0.75/$3.75 itself doubles on 1 January, and DeepSeek just proved a Chinese flagship can quadruple overnight. The durable engineering answer is a model-agnostic router that can shift workloads across providers on price and latency without a rewrite.

4. Price is now one of three axes — watch speed too. Real capability is getting cheaper fast on the American side, and even DeepSeek’s hike leaves it far below Western frontier pricing in absolute terms. But cost is not the only front opening up: in the same week, OpenAI put its flagship on a 750-tokens-per-second “Ultrafast” tier, a sign that latency is becoming a competitive axis alongside intelligence and price. The catch is that “cheapest” now comes with asterisks — surge windows, hard-dated intro rates, unreproducible benchmarks, export-control and government-gating overhang on both sides. Price is finally the main event; it is also more complicated than a single number.

Honest caveats

The verdict

The single most useful thing to take from this week is a deleted assumption: “Chinese model = cheapest” is no longer a rule you can budget on. DeepSeek’s flagship has moved up-market, complete with surge pricing that lands hardest on European working hours, while OpenAI, Anthropic and Google race each other down. For high-volume, cost-sensitive work, the value leaders today are GPT-5.6 Luna, Gemini 3.7 Flash, and DeepSeek V4-Flash off-peak — in that order for most Western teams. For frontier reasoning, GPT-5.6 Sol at $4/$20 is now the cheapest option on paper — with the caveat that it reverts on 21 November, which makes Claude Opus 5 the cheapest frontier rate you can actually plan a year around. But the recommendation that outlasts this month is structural: build so you can switch. See the DeepSeek V4-Pro vs Claude Opus 4.8 breakdown and the best AI coding tools guide for where each model earns its price — and treat the current leader as a tenant, not a landlord.

Update, 21 August 2026: the “build so you can switch” advice got considerably more expensive to take for granted. Stripe has agreed to acquire OpenRouter — the gateway most teams used to make switching cheap — at a reported $7.5 billion, and Ramp launched a free competing router the same day. The mechanics of switching have not changed; who owns them has. See the neutral gateway just got bought and cloned in 24 hours. Ramp’s August AI Index adds the demand-side half of the same picture: premium tiers are struggling to hold their premium — Fable 5 took only 6% of Anthropic’s tokens in July — while the most sophisticated buyers shift volume to open-weight models on third-party serving platforms. Prices are falling because the top of the market is not holding.

Update, 22 August 2026 — two corrections and one reversal. First, the correction: this article listed DeepSeek V4-Flash at $0.14/$0.28 and described it as unchanged. DeepSeek’s live pricing page now shows $0.22/$0.66 off-peak and $0.44/$1.32 at peak; the cheap tier moved too, and the table and prose above have been corrected. From 23 August 2026 the off-peak rate applies across all of Saturday and Sunday (Beijing time), which makes weekend batch work materially cheaper. Second, the reversal: this article’s verdict named Claude Opus 5 the frontier value pick. Eight days later OpenAI cut GPT-5.6 Sol to $4/$20, undercutting Opus 5 on both input and output — but only until 21 November 2026, because the rate is explicitly a three-month promotion. That eight-day turnover is now the story: three of the five prices in the table above carry an expiry date, so “which model is cheapest” has become a question with a shelf life shorter than most procurement cycles. Separately, DeepSeek added vision to V4-Flash at no price premium on 21 August, which resets the floor for multimodal work without changing the text rates above.

Frequently asked questions

Is DeepSeek still the cheapest AI model?

Not for its flagship, and not by the margin people assume. From 17 August 2026 DeepSeek-V4-Pro's peak-hour rate is roughly $1.3 per million input tokens and $3.8 per million output tokens (9 and 27 yuan). That is more expensive than OpenAI's GPT-5.6 Luna at $0.20/$1.20 and Google's Gemini 3.7 Flash at $0.75/$3.75. DeepSeek's cheap tier, V4-Flash, is unchanged at $0.14/$0.28 and remains one of the lowest-priced capable models anywhere — so 'DeepSeek is cheap' is now only true if you specifically use Flash, off-peak, and can live with unreproducible benchmarks. The blanket assumption that any DeepSeek call is the budget option no longer holds.

What is DeepSeek's new peak/off-peak pricing?

DeepSeek is moving V4 to time-of-day surge pricing. Peak hours are 09:00–12:00 and 14:00–18:00 Beijing time; off-peak is everything else, at roughly half the peak rate. At peak, V4-Pro is about 0.3 yuan per million cached-input tokens, 9 yuan uncached input and 27 yuan output. The 'up to 1,100%' headline is the cache-hit input jump specifically (from about 0.025 to 0.3 yuan); output roughly quadruples. The Beijing peak windows map onto European morning and afternoon working hours, so EU teams sit inside the expensive band while US business hours fall almost entirely in off-peak.

Why are OpenAI and Anthropic cutting prices while DeepSeek raises them?

Two different pressures. US labs are competing on cost per unit of useful work because capability leads no longer last a month, and buyers now optimise their inference bill harder than their benchmark scores — so OpenAI cut GPT-5.6 Luna 80%, Anthropic priced Opus 5 at half of Fable 5, and Google halved Flash. DeepSeek is going the other way because its official V4-Pro is a genuinely stronger, more expensive-to-serve flagship, and because serving frontier-scale traffic at near-zero margin is not sustainable — surge pricing is how it protects capacity at peak. In short, the Americans are buying share and the Chinese leader is finally pricing for cost.

Which AI model gives the best value per token right now?

It depends on the job. For high-volume, cost-sensitive work where a mid-tier model is good enough, GPT-5.6 Luna ($0.20/$1.20) and Gemini 3.7 Flash ($0.75/$3.75, through year-end) are the value leaders, with DeepSeek V4-Flash ($0.14/$0.28) cheapest of all if you accept its caveats. For frontier reasoning, **GPT-5.6 Sol took the lead on 21 August at $4/$20**, below Claude Opus 5 ($5/$25) and far below Fable 5 ($10/$50) — but Sol's rate is a promotion that expires 21 November 2026, so Opus 5 remains the cheapest frontier model with no stated expiry date. The honest answer is that the 'cheapest strong model' title now changes hands almost monthly, so the durable move is to keep your stack model-agnostic rather than to chase the current leader.

Should I switch away from DeepSeek because of the price hike?

Switch your assumption, not necessarily your provider. If you were routing bulk traffic to V4-Pro purely because it was cheap, re-run the math against GPT-5.6 Luna, Gemini 3.7 Flash and DeepSeek's own V4-Flash — the flagship is no longer the obvious budget choice, especially during European daytime. If you use V4-Flash off-peak it is still extremely cheap. The broader lesson is to build a router that can move workloads between providers on price and latency, because this is the third repricing of the market in six weeks and it will not be the last.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.