Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Apr 25, 2026
·
modelsopen-sourcebenchmarks

DeepSeek V4 launches: V4 Pro tops LiveCodeBench, V4 Flash undercuts everyone

A year after V3 shocked Silicon Valley, DeepSeek shipped V4 Pro and V4 Flash on April 24, 2026. The headline number is unambiguous: V4-Pro-Max scored 93.5 on LiveCodeBench Pass@1 — higher than any Western model evaluated.

The numbers

V4 Pro ($1.74 / $3.48 per million tokens, 1M context window)

V4 Flash ($0.14 / $0.28 per million tokens)

For comparison: Claude Opus 4.7 scored 87.6% SWE-bench Verified at $5/$25 per million tokens. V4 Pro is a few points behind Opus 4.7 at roughly a third the input price and an eighth the output price. That ratio is what makes DeepSeek genuinely disruptive, not the leaderboard photo.

What’s actually new architecturally

Hybrid Attention Architecture. DeepSeek’s main technical claim — improves long-conversation memory across queries. Combined with the 1M-token context window, V4 Pro can hold an entire codebase or document collection in working memory.

1 trillion total parameters, 32B active. Mixture-of-Experts at frontier scale. The active parameter count keeps inference costs manageable; the total parameter count keeps capability ceiling high.

Huawei Ascend 950PR optimization. V4 is reportedly optimized for Chinese-domestic accelerators. Politically interesting — DeepSeek is decoupling from Nvidia. Practically interesting — if the Ascend optimizations hold, this is the first frontier-class model not bottlenecked by Nvidia GPU supply.

Why this matters for buyers

API users running high-volume: V4 Flash at $0.14/$0.28 makes Claude Haiku 4.5 look expensive and GPT-5.4 Nano look unjustifiable. For batch processing, classification, summarization, or any high-volume workflow where 90% of frontier capability is enough, V4 Flash is now the cost-of-first-resort.

Coding API users: V4 Pro’s LiveCodeBench leadership is genuine. If you’re running Codex-style agents at high volume, V4 Pro at $1.74 input vs Opus 4.7 at $5 input is a 65% cost reduction with marginally lower scores on most coding benchmarks. The break-even calculation favors V4 Pro for almost any volume.

Privacy-sensitive enterprises: Same caveat as V3 — DeepSeek’s hosted API is in China. Data residency, IP exposure, and Chinese government access remain unresolved concerns for regulated industries. The model weights are open, so self-hosting is a path. Several US providers (Together, Fireworks, OpenRouter) are likely to host V4 within the next two weeks.

Solo developers and hobbyists: The DeepSeek consumer chat app is free. If you don’t need ChatGPT’s voice/images/Sora ecosystem, the chat is competitive with the GPT-5.4-tier free experience at zero cost.

What V4 doesn’t change

Frontier still has a moat. DeepSeek’s own line — “3 to 6 months behind GPT-5.4 and Gemini 3.1 Pro” — is. On agentic coding (SWE-Pro, Terminal-Bench), V4 Pro trails. On hard math (FrontierMath Tier 4 where GPT-5.5 just hit 39.6%), DeepSeek hasn’t published competitive numbers. The frontier is real; V4 is closer than V3 was, but it’s not parity.

Geopolitical risk hasn’t gone away. US export controls, potential model bans, supply chain decoupling — none of this is settled. Building a production stack on DeepSeek V4 means accepting that policy risk.

The ecosystem is thinner. No Custom GPTs equivalent. No native voice mode. No first-party image generation. DeepSeek is a model, not a platform. Compared to ChatGPT or Gemini, you’re getting raw capability, not product surface area.

The 12-month picture

A year ago V3 dropped and the entire AI category recalibrated on price. The major closed-source labs responded by either:

  1. Keeping prices flat and shipping better models (Anthropic — Opus 4.7 at $5/$25 is a real value play)
  2. Doubling prices on the new tier and offering smaller cheaper variants (OpenAI — GPT-5.5 just doubled to $5/$30)
  3. Cutting the input price to undercut DeepSeek directly (Google — Gemini 3.1 Pro at $2/M input)

V4 will trigger the second round of this cycle. Watch for Anthropic and OpenAI to either ship cheaper tiers or accept they’re losing the cost-sensitive segment to DeepSeek + Gemini.

What I’d actually do

For most consumer users: Stick with what you have. V4 Flash via the DeepSeek chat app is a solid free option, but ChatGPT/Claude/Gemini ecosystems are stickier than the small marginal cost difference.

For API-heavy automation: Try V4 Flash on your batch workloads. It will probably be 60-80% cheaper for similar quality.

For coding API: V4 Pro is now in the “consider seriously” tier alongside Claude Opus 4.7 and Gemini 3.1 Pro. The geopolitical caveat applies.

For enterprise: The data residency and policy questions still rule out the hosted API for regulated work. Self-hosting V4 weights on your own infrastructure is the path — and now genuinely viable, since the open-weight ecosystem (Llama 4, Mistral Small 4, Qwen 3.6, GLM-5.1, V4) is competitive end-to-end.


Related:

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.