AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Aug 22, 2026
·
deepseekopen-weightspricingbenchmarksagentschina

DeepSeek's cheap tier just beat its own flagship — and a pricing change is coming that hits Europe hardest

TL;DR: DeepSeek shipped V4-Flash-0731 on 31 July 2026 — same 284B-parameter MoE architecture as the April preview (13B active per token, 1M context), with every gain coming from a rebuilt post-training pipeline rather than a new design. On DeepSeek’s own numbers it beats V4-Pro-Preview on all seven published benchmarks — Terminal Bench 2.1 82.7 vs 72.1, DeepSWE 54.4 vs 12.8 — at roughly a third of Pro’s price ($0.14/$0.28 per million vs $0.435/$0.87), under an MIT licence. Two things complicate it. The agent scores were produced with “the minimal mode of DeepSeek Harness, which has not been released,” so they cannot be independently reproduced — the third unverifiable benchmark claim in a month. And DeepSeek’s documentation confirms a peak/off-peak pricing policy at 2× the regular rate during 09:00–12:00 and 14:00–18:00 Beijing time, which lands on European working hours and misses US ones entirely. Update (14 Aug): the policy is now confirmed to take effect 17 August 2026, bundled with the official V4-Pro launch and a broader V4 price rise of up to 1,100% on some line items — full analysis here.

What shipped

DeepSeek released the production build of its Flash model on 31 July 2026, under the name DeepSeek-V4-Flash-0731, replacing the April preview on the same API endpoint.

The interesting part is what did not change. Architecture, parameter count (284B total, 13B activated per token via mixture-of-experts), context window (1M tokens, 384K max output), latency profile and endpoint are all identical to the April preview.

Everything DeepSeek is claiming comes from re-post-training the same base model, with a pipeline rebuilt around coding, tool use, agentic loops and reasoning. This is the April model, taught better.

Caixin reports the release landed “slightly behind its mid-July target” and, notably, arrived without the V4-Pro iteration expected alongside it.

The headline: the cheap tier overtook the expensive one

Here are DeepSeek’s reported figures, with Opus 4.8 as the reference point the company chose:

BenchmarkV4-Flash-0731V4-Flash PreviewV4-Pro PreviewOpus 4.8
Terminal Bench 2.182.761.872.185.0
NL2Repo54.239.438.569.7
Cybergym76.738.752.783.1
DeepSWE54.47.312.858.0
Toolathlon-Verified70.349.755.976.2
Agents’ Last Exam25.215.816.525.7
AutomationBench Public25.110.812.827.2

Two readings, both worth having.

The impressive one: the retrained Flash beats the bigger, more expensive Pro-Preview on every single row. On DeepSWE the gap is not incremental — 54.4 against 12.8 is a different class of model. Post-training, not parameter count, produced that.

The sceptical one: the comparison is against the Pro preview. DeepSeek did not ship a current Pro to compare against, and Caixin notes the absence explicitly. Beating your own older snapshot is a weaker claim than beating your own shipping flagship, and the framing conveniently obscures which one is happening.

There is also a quieter tell. DeepSeek benchmarked against Opus 4.8 — a model Anthropic superseded on 24 July with the Opus 5 launch, a week before this release. Comparing against last month’s frontier is a choice, not an oversight, and it makes the gap look smaller than it is.

Even against the older Opus, Flash trails on all seven. The story is not “DeepSeek caught the frontier.” It is “DeepSeek got much closer to the frontier at a third of its own previous price and a fraction of Anthropic’s.”

The caveat that makes the numbers unverifiable

Buried in DeepSeek’s own materials: the code-agent tasks were run with “the minimal mode of DeepSeek Harness, which has not been released.”

That single sentence removes the ability of anyone outside DeepSeek to reproduce the agent scores. The weights are public and MIT-licensed — you can run the model. You cannot run their scaffolding, and on agentic benchmarks the scaffolding is a large fraction of the score.

This is now a pattern, and it is the third instance in under a month:

None of these is fraud. All three make vendor benchmark claims less useful than they were a year ago. The practical response for a buyer is unchanged and boring: treat launch-day agent scores as marketing until someone independent reproduces them, and weight your own evaluation on your own tasks far more heavily than any published table.

The pricing story is better than the benchmark story

DeepSeek’s published API pricing, verified against its own documentation:

Input (cache miss)Input (cache hit)Output
V4-Flash$0.14 / M$0.0028 / M$0.28 / M
V4-Pro$0.435 / M$0.003625 / M$0.87 / M

Flash is roughly one-third of Pro on both input and output. Set against Opus 5 at $5/$25, Flash output is about 89× cheaper. That is not a comparison of equals — Opus 5 is a materially stronger model — but for high-volume agentic work where you are burning tokens on tool calls and retries rather than on hard reasoning, an 89× multiplier changes which architectures are economically possible.

The cache-hit price is the underrated number. At $0.0028 per million input tokens, repeatedly passing a large fixed context — a codebase, a document set, a system prompt — is close to free. For retrieval-heavy agents that re-read the same corpus on every step, that is the line item that usually decides the bill.

The peak-pricing change, and why it lands on Europe

DeepSeek’s documentation states the API “will soon adopt a peak/off-peak pricing policy. During peak hours, prices will be 2x the regular prices.” Peak is defined as 09:00–12:00 and 14:00–18:00 Beijing time (UTC+8), daily. Update (14 August 2026): DeepSeek confirmed the effective date as 17 August 2026, announced alongside the official launch of V4-Pro and a broader V4 price increase of 50% to 1,100% depending on the line item — the full picture is in our price-war analysis.

Nobody appears to have mapped what that means outside China. Doing it:

RegionPeak windows in local time
Central Europe (CEST, UTC+2)03:00–06:00 and 08:00–12:00
UK (BST, UTC+1)02:00–05:00 and 07:00–11:00
India (IST, UTC+5:30)06:30–09:30 and 11:30–15:30
US Eastern (EDT, UTC−4)21:00–00:00 and 02:00–06:00
US Pacific (PDT, UTC−7)18:00–21:00 and 23:00–03:00

The asymmetry is stark. A European team’s entire morning — 08:00 to 12:00 — sits inside the 2× window. An Indian team loses most of its working day. A US team is untouched: every peak hour falls at night or before dawn.

This is not a conspiracy; it is what happens when a Chinese company prices for its domestic load curve. But the consequence for a European buyer is concrete. If you are running batch agent workloads out of Frankfurt or Dublin, moving them past midday local time — or into the genuinely cheap 12:00–20:00 CEST band — will roughly halve that portion of the bill for zero engineering effort.

It also quietly erodes the headline number. A European shop doing most of its inference during business hours is not paying $0.28 per million output tokens. It is paying closer to $0.56 for a meaningful share of its volume, which narrows the gap to GLM-5.2 and other open-weight alternatives.

Self-hosting: permissive licence, unforgiving hardware

The weights are MIT-licensed and ungated — more permissive than most open-weight releases, including several that call themselves open. No acceptable-use rider, no gated download, commercial deployment allowed.

The hardware is the wall: roughly 110 GB at 3-bit quantisation, or a full 4×GB300 node at full precision. That puts genuine self-deployment in the same bracket as Kimi K3 — the licence says you are free, the memory requirement says you are renting. For most teams “open weights” here means insurance against API changes and the option to audit, not a plan to actually run it.

Why this matters

Post-training is where the gains are now. Same architecture, same parameter count, dramatically different capability. That is the clearest single data point of 2026 for the argument that the frontier has shifted from scaling runs to training recipes — and it is the reason a vendor’s cheap tier can eat its own flagship inside four months.

It compresses the value tier further. DeepSeek, Kimi K3, and GLM are now clustered in a band that did not exist eighteen months ago: agent-capable, open-weight, and priced where you stop thinking about the bill. The pressure that produced DeepSeek’s permanent 75% Pro cut in May has not let up.

The verification gap is becoming the story. When three consecutive major releases ship numbers nobody outside the vendor can reproduce, the benchmark table stops being evidence and becomes a press release with a monospace font.

The geopolitical overhang has not gone away. US policy is actively moving against Chinese open-weight models — see the proposed open-source curbs and chip provisions. Building critical infrastructure on a Chinese API carries a policy risk that has nothing to do with model quality.

Honest caveats

Every benchmark here is DeepSeek’s own, and the agent subset is unreproducible by design. No independent evaluation had been published at the time of writing.

The Pro comparison is against a preview, not a current release, and DeepSeek did not ship a Pro update alongside it.

Reports on pricing conflict. Caixin describes the update as “further reducing API costs” with “up to 50% cheaper” costs; other coverage says pricing is unchanged from the April preview. The figures in this article come from DeepSeek’s own live pricing page and are what you will actually be charged today — but the peak/off-peak policy has no announced start date, so today’s price is not a commitment.

Caixin reports Flash still trails Moonshot’s Kimi K3. No head-to-head using an identical harness exists, so treat the ordering between them as unsettled.

Data residency is unchanged. API traffic goes to Chinese servers. For anything you would not send to a Chinese jurisdiction, the MIT licence is the relevant feature, not the API price — and the hardware bill above is what that actually costs.

What to do

If you already use DeepSeek: nothing breaks — same endpoint, same price today. If you are in Europe or India, look at when your batch jobs run and shift what you can out of the peak windows before the policy takes effect.

If you are choosing a value-tier agent model: Flash is the strongest price-per-capability option on paper right now, with the caveat that “on paper” is doing real work in that sentence. Run your own tasks before committing; the published agent scores cannot be checked.

If you are on Claude Code or a frontier model for agent work: this does not displace it. Flash trails even the superseded Opus 4.8 across the board. What it changes is the calculus for the high-volume, lower-stakes half of an agent workload — the retries, the tool-call chatter, the retrieval passes — where paying frontier prices was always hard to justify.

Full context in the DeepSeek review, updated with these figures.

Update, 22 August 2026 — the peak pricing landed, and V4-Flash can now see. The peak-hours policy flagged above as “effective date not yet announced” took effect 17 August 2026 and is now on DeepSeek’s published pricing page: V4-Flash runs $0.22/$0.66 off-peak and $0.44/$1.32 at peak per million tokens, with peak windows at 01:00–04:00 and 06:00–10:00 UTC — squarely across European mornings, as predicted. One improvement: from 23 August 2026 the cheaper off-peak rate applies across all of Saturday and Sunday (Beijing time), which makes weekend batch work materially cheaper. Separately, on 21 August DeepSeek added image understanding to V4-Flash at no price premium via deepseek-v4-flash-vision-exp, with images capped at 384 tokens each. The unreleased-harness caveat this article raised about V4-Flash-0731’s agent scores applies unchanged to the vision model’s benchmarks. For the wider picture, see the frontier price crown’s new expiry date.


Related: The EU AI Act’s transparency rules are live · July 2026 in AI — what changed for buyers · DeepSeek V4 launch, April 2026

Frequently asked questions

Is DeepSeek V4-Flash-0731 actually better than V4-Pro?

On the benchmarks DeepSeek published, the retrained Flash beats V4-Pro-Preview on all seven — including Terminal Bench 2.1 (82.7 vs 72.1) and DeepSWE (54.4 vs 12.8). The comparison is against the Pro *preview*, not a current Pro release, and DeepSeek did not ship a V4-Pro update alongside this one. So the honest framing is that the cheap tier has overtaken an older snapshot of the expensive tier, not that Pro has been retired.

What does V4-Flash cost?

Per DeepSeek's own pricing page: $0.14 per million input tokens on a cache miss, $0.0028 on a cache hit, and $0.28 per million output tokens. V4-Pro is $0.435 input / $0.87 output. That makes Flash roughly one-third the price of Pro on both sides, and still the cheapest credible agent-capable model on the market.

What is the peak-hours pricing change?

DeepSeek's documentation states the API will adopt peak/off-peak pricing, with prices 2x the regular rate during 09:00–12:00 and 14:00–18:00 Beijing time daily. Update: on 14 August 2026 DeepSeek confirmed the policy takes effect 17 August, announced alongside the official V4-Pro launch and a broader V4 price increase of 50% to 1,100% on some line items — see our [full analysis of the flip](/news/ai-price-war-flips-deepseek-raises-us-labs-cut-2026-08-14/). Those windows map to roughly 03:00–06:00 and 08:00–12:00 CEST, meaning European morning working hours fall inside the expensive window while US business hours fall entirely outside it.

Can I verify DeepSeek's benchmark claims myself?

Not fully. DeepSeek states the code-agent tasks were run using the minimal mode of DeepSeek Harness, which has not been released. Without the harness, the agent numbers cannot be reproduced. The model weights themselves are MIT-licensed and ungated, so the model is verifiable even though the scores are not.

Can I self-host it?

Yes — the weights are MIT-licensed and ungated, which is more permissive than most open-weight releases. The hardware is the constraint: roughly 110 GB at 3-bit quantisation, or a full 4×GB300 node for full precision. That puts single-GPU self-hosting out of reach and makes the API the realistic option for most teams.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.