The AI price war just flipped: DeepSeek is raising prices up to 1,100% while OpenAI and Anthropic cut
TL;DR: The AI price war just inverted. On 13–14 August 2026 DeepSeek shipped the official build of its flagship V4-Pro and, the same day, said it is raising API prices by 50% to 1,100% with peak/off-peak surge pricing from 17 August. At peak, V4-Pro runs about 9 yuan (≈$1.3) input and 27 yuan (≈$3.8) output per million tokens — roughly a 4× jump on output — while the “1,100%” figure is the cache-hit input rate specifically. In the same window, US labs cut: OpenAI trimmed GPT-5.6 Luna 80% to $0.20/$1.20, Anthropic launched Claude Opus 5 at $5/$25 — half of Fable 5’s $10/$50 — and Google put Gemini 3.7 Flash out at $0.75/$3.75, half of 3.6 Flash. Financial Times data shows US model prices have fallen materially since mid-July, narrowing the US-vs-China gap to roughly 19% on input, 14% on output. The takeaway for buyers: the reflex that “DeepSeek = cheapest” is now wrong for its flagship — GPT-5.6 Luna and Gemini Flash undercut V4-Pro, and only DeepSeek’s V4-Flash ($0.14/$0.28) is still rock-bottom. The durable move is a model-agnostic router, because the “cheapest strong model” title now changes hands monthly.
What just happened
For most of the last two years, AI budgeting ran on one assumption: American frontier models were the expensive, best-in-class option, and Chinese open-weights models — DeepSeek above all — were the cheap fallback you reached for when the bill got scary. In the space of about a week in mid-August 2026, that assumption broke in both directions at once.
On 13–14 August, DeepSeek released the production build of its flagship, DeepSeek-V4-Pro, and simultaneously announced a sweeping change to its API pricing. Per Caixin Global, which first reported the change, rates are rising by 50% to as much as 1,100% depending on the model, the token type and the time of day, and the new prices take effect 17 August 2026. The most-quoted “1,100%” number is the increase on cache-hit input tokens — the cheapest line item — which jumps from roughly 0.025 yuan to about 0.3 yuan per million. The headline workloads move less dramatically but still sharply: at peak, uncached input runs about 9 yuan (≈$1.3) and output about 27 yuan (≈$3.8) per million tokens, roughly quadrupling V4-Pro’s previous $0.435/$0.87. DeepSeek is also introducing peak/off-peak surge pricing — peak windows of 09:00–12:00 and 14:00–18:00 Beijing time, with off-peak roughly half price.
At the same time, the American labs kept moving the other way:
- OpenAI cut its cheapest tier, GPT-5.6 Luna, by 80% in late July, to $0.20 input / $1.20 output per million, and trimmed the mid-tier Terra ~20% to $2/$12.
- Anthropic launched Claude Opus 5 at $5/$25 per million — half the price of its flagship Fable 5 ($10/$50) — and, per reporting, shelved a planned September increase on Claude Sonnet 5, holding it at its $2/$10 introductory rate.
- Google shipped Gemini 3.7 Flash on 13 August at $0.75/$3.75 — half of 3.6 Flash — through year-end.
Financial Times data cited across the coverage shows the prices customers actually pay for leading US models have declined materially since mid-July, and the gap between US closed-source pricing and the Chinese market has narrowed to roughly 19% on input and 14% on output. Put the two trends together and you get the story: the cheap option is getting more expensive, and the expensive option is getting cheaper.
The new pricing map
Here is where the major API models sit as of mid-August 2026, per each vendor’s published pricing and the reporting above. All figures are per million tokens, input / output, in USD; DeepSeek’s new rates are converted from yuan and are approximate.
| Model | Tier | Input | Output | Note |
|---|---|---|---|---|
| Anthropic Fable 5 | Frontier | $10 | $50 | Most expensive frontier flagship |
| OpenAI GPT-5.6 Sol | Frontier | $4 | $20 | Cut 21 Aug; promo rate expires 21 Nov 2026 |
| Anthropic Claude Opus 5 | Frontier | $5 | $25 | Launched 24 Jul at half of Fable 5 |
| Moonshot Kimi K3 | Large open-weight | — | ~$15 | Cheapest “large frontier-class” option |
| OpenAI GPT-5.6 Terra | Mid | $2 | $12 | Cut ~20% in late July |
| Anthropic Claude Sonnet 5 | Mid | $2 | $10 | Intro rate; Sept rise reportedly scrapped |
| DeepSeek V4-Pro (from 17 Aug, peak) | Flagship | ≈$1.3 | ≈$3.8 | Was $0.44/$0.87; off-peak ≈ half |
| Google Gemini 3.7 Flash | Workhorse | $0.75 | $3.75 | Intro through 31 Dec, then $1.50/$7.50 |
| OpenAI GPT-5.6 Luna | Cheap | $0.20 | $1.20 | Cut 80% in late July |
| DeepSeek V4-Flash | Cheap | $0.22 | $0.66 | Off-peak; doubles at peak. Corrected 22 Aug |
Two things jump out. First, DeepSeek’s flagship is no longer a budget product: at peak, V4-Pro’s ≈$3.8 output lands level with Gemini 3.7 Flash’s $3.75 and more than triples GPT-5.6 Luna’s $1.20 — the two American tiers buyers actually reach for on price. The only DeepSeek line that still sits near the price floor is V4-Flash, though DeepSeek’s live pricing page now puts it at $0.22/$0.66 off-peak and $0.44/$1.32 at peak — the $0.14/$0.28 figure we published on 14 August is no longer what the vendor charges. Second, the American mid-and-cheap tiers — Luna, Terra, Sonnet 5, Gemini Flash — now occupy the value band that Chinese models owned a year ago.
Why the flip is happening
The two directions have different causes, and both are rational.
The US labs are buying share with price because capability leads have stopped lasting. When a frontier lead evaporates within weeks — Grok 4.6 matching GPT-5.6 on the intelligence index at launch, Z.ai’s GLM-5.3 and DeepSeek resetting the open-weights floor monthly — “we have the smartest model” is no longer a durable pitch. So the labs compete on the number buyers can actually control: cost per unit of useful work. Anthropic’s Opus 5 is the cleanest example — it did not lower the sticker of its top tier, it shipped more capability at the Opus price band while halving the flagship rate. OpenAI’s Luna cut and Google’s Flash discount are the same move aimed at the highest-volume workload in AI: coding agents that burn tokens by the million.
DeepSeek is raising prices because its flagship finally costs money to run — and because near-zero margin is not a business. The official V4-Pro is a stronger, more expensive-to-serve model than the preview it replaces. DeepSeek’s answer is not a flat hike but surge pricing: charge more when GPUs are saturated (Chinese business hours) and less when they are idle. That is a capacity-management tool as much as a revenue one, and it signals that even the poster child of “too cheap to meter” AI has hit the same infrastructure wall everyone else has. The cheap tier moving far less than the flagship is the tell: DeepSeek is segmenting, pushing price-sensitive traffic to Flash and monetising performance on Pro.
Why this matters
1. The “DeepSeek is cheapest” reflex is now a budgeting bug. If you route bulk traffic to a DeepSeek flagship endpoint on the assumption that it is automatically the low-cost option, you are, from 17 August, very possibly overpaying versus GPT-5.6 Luna or Gemini 3.7 Flash — especially during European daytime, which sits squarely in DeepSeek’s peak window. Re-run the math per workload; the map above is the starting point, your own token mix is the answer.
2. Surge pricing punishes geography. Time-of-day pricing is new to this market, and it is not neutral. DeepSeek’s peak hours (09:00–12:00 and 14:00–18:00 Beijing) overlap European mornings and afternoons almost perfectly, while US business hours fall in the cheap off-peak band. If you are an EU team defaulting to DeepSeek for cost, you are now paying peak rates during exactly the hours you work. This is the same policy we flagged as pending on V4-Flash in early August — it is now confirmed and dated.
3. Cheap is no longer a moat, so lock-in is the real risk. The uncomfortable truth of this market is that the value crown rotates monthly: Luna in late July, Gemini Flash in mid-August, whatever ships next month. Optimising your architecture around this week’s cheapest strong model is how you get stranded when it re-prices — and note that Gemini Flash’s $0.75/$3.75 itself doubles on 1 January, and DeepSeek just proved a Chinese flagship can quadruple overnight. The durable engineering answer is a model-agnostic router that can shift workloads across providers on price and latency without a rewrite.
4. Price is now one of three axes — watch speed too. Real capability is getting cheaper fast on the American side, and even DeepSeek’s hike leaves it far below Western frontier pricing in absolute terms. But cost is not the only front opening up: in the same week, OpenAI put its flagship on a 750-tokens-per-second “Ultrafast” tier, a sign that latency is becoming a competitive axis alongside intelligence and price. The catch is that “cheapest” now comes with asterisks — surge windows, hard-dated intro rates, unreproducible benchmarks, export-control and government-gating overhang on both sides. Price is finally the main event; it is also more complicated than a single number.
Honest caveats
- DeepSeek’s USD figures are converted from yuan and are approximate. The primary numbers are DeepSeek’s published yuan rates (≈0.3 / 9 / 27 yuan per million for cached input / uncached input / output at peak) as reported by Caixin; dollar equivalents move with the exchange rate. Treat them as “roughly,” not to the cent.
- The “1,100%” is one line item, not the whole model. It is the cache-hit input increase. Output “only” roughly quadruples. Both are large; neither means every V4-Pro call costs 12× more.
- Sonnet 5’s cancelled increase is from reporting, not an Anthropic price page we can point to today. The $2/$10 introductory rate is confirmed; the scrapped September rise is second-hand and should be treated as such.
- Benchmarks are not in scope here, and DeepSeek’s remain hard to reproduce — the agentic scores rely on a harness the company has not released, a caveat that predates this pricing change and still applies.
- Prices change faster than articles. This map is a mid-August 2026 snapshot. Before you commit spend, check each vendor’s live pricing page; three of these numbers have a built-in expiry date already.
The verdict
The single most useful thing to take from this week is a deleted assumption: “Chinese model = cheapest” is no longer a rule you can budget on. DeepSeek’s flagship has moved up-market, complete with surge pricing that lands hardest on European working hours, while OpenAI, Anthropic and Google race each other down. For high-volume, cost-sensitive work, the value leaders today are GPT-5.6 Luna, Gemini 3.7 Flash, and DeepSeek V4-Flash off-peak — in that order for most Western teams. For frontier reasoning, GPT-5.6 Sol at $4/$20 is now the cheapest option on paper — with the caveat that it reverts on 21 November, which makes Claude Opus 5 the cheapest frontier rate you can actually plan a year around. But the recommendation that outlasts this month is structural: build so you can switch. See the DeepSeek V4-Pro vs Claude Opus 4.8 breakdown and the best AI coding tools guide for where each model earns its price — and treat the current leader as a tenant, not a landlord.
Update, 21 August 2026: the “build so you can switch” advice got considerably more expensive to take for granted. Stripe has agreed to acquire OpenRouter — the gateway most teams used to make switching cheap — at a reported $7.5 billion, and Ramp launched a free competing router the same day. The mechanics of switching have not changed; who owns them has. See the neutral gateway just got bought and cloned in 24 hours. Ramp’s August AI Index adds the demand-side half of the same picture: premium tiers are struggling to hold their premium — Fable 5 took only 6% of Anthropic’s tokens in July — while the most sophisticated buyers shift volume to open-weight models on third-party serving platforms. Prices are falling because the top of the market is not holding.
Update, 22 August 2026 — two corrections and one reversal. First, the correction: this article listed DeepSeek V4-Flash at $0.14/$0.28 and described it as unchanged. DeepSeek’s live pricing page now shows $0.22/$0.66 off-peak and $0.44/$1.32 at peak; the cheap tier moved too, and the table and prose above have been corrected. From 23 August 2026 the off-peak rate applies across all of Saturday and Sunday (Beijing time), which makes weekend batch work materially cheaper. Second, the reversal: this article’s verdict named Claude Opus 5 the frontier value pick. Eight days later OpenAI cut GPT-5.6 Sol to $4/$20, undercutting Opus 5 on both input and output — but only until 21 November 2026, because the rate is explicitly a three-month promotion. That eight-day turnover is now the story: three of the five prices in the table above carry an expiry date, so “which model is cheapest” has become a question with a shelf life shorter than most procurement cycles. Separately, DeepSeek added vision to V4-Flash at no price premium on 21 August, which resets the floor for multimodal work without changing the text rates above.
Frequently asked questions
Is DeepSeek still the cheapest AI model?
Not for its flagship, and not by the margin people assume. From 17 August 2026 DeepSeek-V4-Pro's peak-hour rate is roughly $1.3 per million input tokens and $3.8 per million output tokens (9 and 27 yuan). That is more expensive than OpenAI's GPT-5.6 Luna at $0.20/$1.20 and Google's Gemini 3.7 Flash at $0.75/$3.75. DeepSeek's cheap tier, V4-Flash, is unchanged at $0.14/$0.28 and remains one of the lowest-priced capable models anywhere — so 'DeepSeek is cheap' is now only true if you specifically use Flash, off-peak, and can live with unreproducible benchmarks. The blanket assumption that any DeepSeek call is the budget option no longer holds.
What is DeepSeek's new peak/off-peak pricing?
DeepSeek is moving V4 to time-of-day surge pricing. Peak hours are 09:00–12:00 and 14:00–18:00 Beijing time; off-peak is everything else, at roughly half the peak rate. At peak, V4-Pro is about 0.3 yuan per million cached-input tokens, 9 yuan uncached input and 27 yuan output. The 'up to 1,100%' headline is the cache-hit input jump specifically (from about 0.025 to 0.3 yuan); output roughly quadruples. The Beijing peak windows map onto European morning and afternoon working hours, so EU teams sit inside the expensive band while US business hours fall almost entirely in off-peak.
Why are OpenAI and Anthropic cutting prices while DeepSeek raises them?
Two different pressures. US labs are competing on cost per unit of useful work because capability leads no longer last a month, and buyers now optimise their inference bill harder than their benchmark scores — so OpenAI cut GPT-5.6 Luna 80%, Anthropic priced Opus 5 at half of Fable 5, and Google halved Flash. DeepSeek is going the other way because its official V4-Pro is a genuinely stronger, more expensive-to-serve flagship, and because serving frontier-scale traffic at near-zero margin is not sustainable — surge pricing is how it protects capacity at peak. In short, the Americans are buying share and the Chinese leader is finally pricing for cost.
Which AI model gives the best value per token right now?
It depends on the job. For high-volume, cost-sensitive work where a mid-tier model is good enough, GPT-5.6 Luna ($0.20/$1.20) and Gemini 3.7 Flash ($0.75/$3.75, through year-end) are the value leaders, with DeepSeek V4-Flash ($0.14/$0.28) cheapest of all if you accept its caveats. For frontier reasoning, **GPT-5.6 Sol took the lead on 21 August at $4/$20**, below Claude Opus 5 ($5/$25) and far below Fable 5 ($10/$50) — but Sol's rate is a promotion that expires 21 November 2026, so Opus 5 remains the cheapest frontier model with no stated expiry date. The honest answer is that the 'cheapest strong model' title now changes hands almost monthly, so the durable move is to keep your stack model-agnostic rather than to chase the current leader.
Should I switch away from DeepSeek because of the price hike?
Switch your assumption, not necessarily your provider. If you were routing bulk traffic to V4-Pro purely because it was cheap, re-run the math against GPT-5.6 Luna, Gemini 3.7 Flash and DeepSeek's own V4-Flash — the flagship is no longer the obvious budget choice, especially during European daytime. If you use V4-Flash off-peak it is still extremely cheap. The broader lesson is to build a router that can move workloads between providers on price and latency, because this is the third repricing of the market in six weeks and it will not be the last.
Sources
- Caixin Global — DeepSeek Launches V4-Pro and Raises API Prices by as Much as 1,100%
- Engadget — DeepSeek's AI models are about to cost four times more
- Quartz — DeepSeek raising API prices by up to 1,100% starting Aug. 16
- BigGo Finance — OpenAI and Anthropic Slash Prices as Chinese AI Rivals Flip the Script on Costs
- Tech Startups — Top Tech News Today, August 14, 2026 (DeepSeek, OpenAI, Anthropic, Google pricing)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.