DeepSeek slashes V4-Pro prices 75% — now permanent (was promo through May 5)
Update (May 22-25, 2026): DeepSeek confirmed the 75% cut is now permanent — the discounted rates ($0.435/M input, $0.87/M output, $0.03625/M cache hit) will not roll back after the originally-planned May 31 expiry. InfoWorld characterized the move as the most aggressive frontier-tier price cut of 2026: unlike earlier 2026 budget-tier cuts, this one targets the frontier-capability band where OpenAI, Anthropic, and Google compete. Even at full price V4-Pro already undercut GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro per-token; permanent discount makes the gap structural. Direct downward pressure on Western frontier-lab pricing through H2 2026.
Eight days after V4 Pro shipped, DeepSeek announced a 75% promotional discount on V4-Pro originally running through May 5, 2026, 15:59 UTC (later extended to May 31, then made permanent per the May 22-25 update above). The math is brutal:
| Component | Standard | Promo (-75%) | Western frontier (Opus 4.7) |
|---|---|---|---|
| Input (cache miss) | $1.74/M | $0.435/M | $5.00/M |
| Input (cache hit) | $0.145/M | $0.03625/M | — |
| Output | $3.48/M | $0.87/M | $25.00/M |
For the duration of the promo, V4-Pro is roughly 11x cheaper than Claude Opus 4.7 on input and 29x cheaper on output. Even off-promo, V4-Pro was already 5–7x cheaper.
The 90% cache-hit cut is the bigger story
Buried under the V4-Pro headline: DeepSeek simultaneously reduced input cache-hit charges by 90% across its entire API portfolio (V3.2 and R1 included), permanent — not promotional. For workloads with repeated context (long-running agents, codebase-aware tools, document analysis pipelines), the effective cost is now genuinely a fraction of what it was a week ago.
V3.2 cache-hit input was already the cheapest credible LLM pricing on the market at $0.028/M; that just dropped further.
What this means
For developers building on DeepSeek: if you can hit the May 5 deadline, lock in V4-Pro evaluations now. The promotional pricing makes the per-call economics dramatically better than anything Western for the next three days.
For Western frontier providers: the gap was already uncomfortable; the promo tier compounds the pressure. Anthropic’s $40B Google deal and OpenAI’s IPO-track positioning suggest the response is going to be capacity + ecosystem rather than matching DeepSeek on price.
For DeepSeek review readers: the trade-offs covered there (Chinese content moderation, data-residency concerns for sensitive workloads) are unchanged. The 75% cut doesn’t make DeepSeek a better fit for politically-sensitive research; it just makes the existing cost-performance argument more aggressive for everything else.
The strategic read
A 75% cut applied within eight days of launch isn’t normal pricing behavior. It’s market-share warfare, and the timing — overlapping with GPT-5.5 rollout and the Pentagon AI deals that excluded Anthropic — suggests DeepSeek wants enterprise developers running production benchmarks on V4-Pro this week, before the migration window narrows.
The promo ends May 5. After that, V4-Pro returns to standard pricing — still cheap, but the gap to Western frontier models narrows back from “obvious mispricing” to “just very cheap.”
If you’re cost-sensitive and haven’t benchmarked V4-Pro yet, this is the cheapest week to do it.
Background: how DeepSeek got here
The DeepSeek pricing war isn’t new. V3 launched in December 2024 at API prices roughly 20x lower than GPT-4-class models. R1 launched in early 2025 with reasoning capability comparable to OpenAI o1 at a fraction of the cost. By April 2026, V4 Pro topped LiveCodeBench. The trajectory has been consistent: DeepSeek ships frontier-class quality, prices it dramatically below US competitors, and absorbs the market-share consequences.
The 75% promotional cut applied within eight days of V4-Pro’s launch is the latest move in this strategy. It’s market-share warfare, timed for maximum effect against GPT-5.5 rollouts and US enterprise procurement decisions happening this quarter.
What developers should actually do this week
For teams running LLM-heavy production workloads, the three actions worth taking before May 5:
-
Run benchmarks on V4-Pro at promo pricing. If V4-Pro is competitive on your evaluation set at $0.435/M input, the post-promo price ($1.74/M) is still cheaper than every Western alternative. The benchmarks you collect this week inform your post-promo decision.
-
Audit your context-caching architecture. The permanent 90% cut on cache-hit charges is the bigger long-term story. Workloads that re-send the same context (codebase analysis, document Q&A, agent loops) will benefit dramatically. Refactor your prompts to maximize cache hits; the savings are immediate and permanent.
-
Update your model-routing logic. If you’re running a multi-model architecture (cheap model for routine queries, expensive model for hard ones), re-evaluate the routing thresholds with V4-Pro’s new pricing. Tasks you previously routed to GPT-5.4 because Opus was too expensive may now route to V4-Pro at lower cost than either.
What this means for the broader market
Anthropic, OpenAI, and Google can’t easily match these prices without burning meaningful capital. Anthropic’s recent $40B Google deal provides some runway; OpenAI’s IPO-track positioning makes price wars expensive optically. The expected response is feature differentiation (capacity, safety guardrails, ecosystem maturity) rather than price matching.
For end users on consumer subscriptions (ChatGPT Plus, Claude Pro), the pricing shifts don’t change the immediate experience. The impact is on developers building products on top of LLM APIs — for them, the economics of “what’s possible to ship” just changed materially.
The data residency caveat hasn’t changed
Worth restating: DeepSeek’s hosted API runs on Chinese infrastructure. The 75% cut doesn’t affect the data-residency considerations that rule out DeepSeek for regulated industries, US-government-adjacent work, or anywhere that data sovereignty matters. For those workloads, the path is self-hosting V4-Pro on your own infrastructure (the open weights make this viable), or sticking with Western providers despite the price gap.
For non-sensitive workloads, the May 5 deadline is the action-forcing moment. Three days of cheap benchmarks could change your 2026 LLM stack.
Sources
- DeepSeek V4-Pro 75% Discount (DeepSeek API Docs)
- DeepSeek cuts V4-Pro prices by 75% (TheNextWeb)
- DeepSeek cuts AI model prices by 75% in push against US rivals (Investing.com)
- DeepSeek Slashes V4-Pro API Pricing With Major Discount (Dataconomy)
- DeepSeek's steep V4-Pro price cut escalates AI pricing war (InfoWorld)
- DeepSeek Locks In 75 Percent Price Cut for New V4-Pro AI Model (Winbuzzer)
- DeepSeek V4-Pro 75% Price Cut Is Now Permanent (apidog)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.