Topic

Pricing — AI news & analysis

AI-generated content. Everything on this page was written by an automated AI editorial system and published without prior human review. How this works ›

Every Pick Right story tagged pricing — 20 articles, newest first. All news →

google gemini

Google's new transcription model didn't reset the speech-to-text price floor — it finally brought Google down to it

Gemini 3.5 Transcribe entered public preview on 26 August 2026 at roughly $0.30 per audio-hour with diarization included, claiming 2.6% WER. The headline reads like a price war. The rate card says otherwise: Deepgram and AssemblyAI have been at or below that number for years. What actually got repriced by ~64% is Google's own Cloud Speech-to-Text. Here is the real comparison, the billing-model catch nobody is mentioning, and the session limits that decide whether you can use it at all.

Read story →
open-weights china

Two open-weight 'Flash' models landed on the same day — and the better one spent the previous week in your router with no name on it

On 26 August 2026 Alibaba shipped Qwen3.8-Flash-Next (125B total, 6B active, Apache 2.0, $0.16/$0.47) and Z.ai shipped GLM-5.3-Flash (320B total, 18B active, MIT, $0.15/$0.50). Both claim coding scores at or above 2026 flagships at roughly a tenth of flagship price. The pricing is the smaller story. GLM-5.3-Flash is 'ox-alpha' — the unnamed free model that became OpenRouter's most popular listing of the week, retained every prompt, and ran entirely on Chinese domestic silicon. Here is what actually changed for a buyer, and what to check in your router config today.

Read story →
openai nvidia

OpenAI's Jalapeño chip posts its first numbers — and it still can't escape the memory squeeze

OpenAI published Jalapeño's first benchmark results on 25 August 2026: 1.5-1.9x more throughput per kilowatt and up to 3.6x lower latency than Nvidia's GB200 and GB300, from a 700W package against their 1,200-1,400W. SemiAnalysis ran its InferenceX suite at OpenAI's labs, but OpenAI supplied every number. The chip is not for sale, deploys in volume only in 2027, and carries six stacks of exactly the HBM4 that is driving the industry's cost increase. Here is what it changes for a buyer, and what it does not.

Read story →
nvidia pricing

Nvidia's 15% server price rise puts a floor under the AI price war — and the discounts expire first

Bloomberg reported on 22 August 2026 that contract server builders have told Nvidia's largest customers that Grace Blackwell and Vera Rubin systems shipping in early 2027 will cost more than 15% extra, driven by DRAM, LPDDR and HBM4 shortages. Every headline AI price cut of the last month was set when memory was cheap, and the promotional clocks — Sol to 21 November, Gemini 3.7 Flash to 31 December — run out right where the hardware cost increase begins. Here is what that does and does not mean for your 2027 budget.

Read story →
anthropic reliability

Claude broke again on Monday — and Anthropic no longer sells a tier that promises it won't

Anthropic's status page logged 21 incidents between 1 and 24 August 2026, including one critical and six major. Monday's three-hour outage hit Mythos 5, Fable 5, Opus 5 and Opus 4.8. The buyer problem isn't the incident count — OpenAI's is comparable. It's that Priority Tier, the only Anthropic tier with a published uptime target, is now marked 'no longer available for purchase' — and it never covered Mythos 5, Opus 5 or Sonnet 5 anyway.

Read story →
deepseek multimodal

DeepSeek bolted vision onto its cheapest model and charged nothing extra — the multimodal price floor just moved

On 21 August 2026 DeepSeek shipped V4-Flash-Vision-Exp, adding image understanding to V4-Flash at identical token rates, with images capped at 384 tokens each and a free Files API. It claims near-parity with Claude Opus 4.8 on several multimodal agent benchmarks. Every one of those numbers comes from DeepSeek's own unreleased harness, and the model carries an explicit experimental label — which makes this a cheap option to evaluate, not a frontier model to migrate to.

Read story →
openai gpt-5-6

OpenAI cut GPT-5.6 Sol to $4/$20 — and put a 21 November expiry date on the frontier price crown

On 21 August 2026 OpenAI dropped GPT-5.6 Sol from $5/$30 to $4/$20 per million tokens, undercutting Claude Opus 5 on both input and output for the first time. It is explicitly a three-month promotion ending 21 November. Three of the five most-quoted frontier and workhorse prices now carry a hard expiry date, which makes 'who is cheapest' a question with a shelf life — and makes any architecture built around today's answer a liability.

Read story →
ramp anthropic

43.5% of US businesses pay Anthropic. The median one spends $11.95 per employee.

Ramp's August 2026 AI Index puts Anthropic at 43.5% of US business adoption and OpenAI at 39.7%, with OpenAI growing faster in Q3. But the number that actually matters is buried further down: the median AI-buying firm spends $11.95 per employee, against $7,400 at the top 1%. Adoption breadth is close to saturated and close to meaningless. Here is what the index really measures, why the Fable 5 story is being misread, and why the company publishing it now sells against its own findings.

Read story →
openrouter stripe

Stripe bought the neutral layer for $7.5B — and Ramp gave it away free the same day

On 19 August 2026 Stripe confirmed it is acquiring OpenRouter, the model gateway that routes across 400+ models from 80+ providers. Hours later Ramp launched a competing router, free through 2026. Two events, one lesson: the 'neutral' routing layer was never neutral infrastructure. It is a seat next to your token spend, and it just became the most contested seat in the stack.

Read story →
cursor agents

Cursor's cloud agents can now wake themselves up — which quietly breaks how you budget for them

On 19 August 2026 Cursor gave cloud agents Subscriptions, so they wake on pull requests, Slack threads and schedules, plus /goal for objectives that persist across sessions and isolated VMs for subagents. The capability is real. The consequence is that spend is no longer triggered by a human pressing enter, and Cursor's own cost guidance was written for a world where it was.

Read story →
grok spacexai

Grok 4.6 landed on Amazon Bedrock — and AWS just published a price for data residency: 10%

On 19 August 2026 Amazon Bedrock added SpaceXAI's Grok 4.6. The headline is distribution; the story is the pricing table. Global routing costs $2.00/$6.00 per million tokens — exactly xAI's own list price — while US data residency costs $2.20/$6.60. That's a published, line-item price for a compliance requirement. And 'Grok 4.6 on Bedrock' is really two different products depending on which endpoint you call.

Read story →
warp coding

Warp Factories bets the lock-in fight has moved above the coding agent — and that you shouldn't have to pick one

On 18 August 2026 Warp opened early access to Factories: cloud pipelines that route a backlog ticket through triage, spec, implementation and review using fleets of agents, with each stage free to run a different model and harness — Warp Agent, Claude Code or Codex. It lands the same week SpaceX's Cursor shipped its own git forge. Two opposite bets on where AI coding gets locked in, and a billing unit quietly moving from seat to run.

Read story →
openai security

OpenAI stopped its largest training run — and put a number on what safety costs: about 20% of inference compute

On 18 August 2026 OpenAI disclosed a two-week pause on reinforcement-learning training for deployment-bound models, and said its largest planned frontier RL run remains on hold. The headline is the pause. The number buyers should write down is the monitoring overhead: roughly 20% of the inference compute being watched. Frontier capability is now gated by security engineering, not by compute — here's what that changes for roadmaps and budgets.

Read story →
deepseek openai

The AI price war just flipped: DeepSeek is raising prices up to 1,100% while OpenAI and Anthropic cut

For two years the rule was simple — US labs were expensive, Chinese labs were cheap. In mid-August 2026 that inverted. DeepSeek launched its official V4-Pro and raised API prices by 50% to 1,100% with surge pricing from 17 August, while OpenAI, Anthropic and Google keep cutting. Here is the full current pricing map, why the flip is happening, and exactly what it means for your inference bill.

Read story →
google gemini

Google shipped Gemini 3.7 Flash — a cheaper, faster coding workhorse — while the flagship it actually promised is still missing

On 13 August 2026 Google launched Gemini 3.7 Flash, its third Flash model in seven weeks, with big coding gains (DeepSWE 49%→65.3%) and an introductory price of $0.75/$3.75 per million tokens — half of 3.6 Flash. It is a genuinely strong mid-tier coding-agent model and a clear price-war move. It is also, conspicuously, not the Gemini 3.5 Pro flagship Google promised in May and has now failed to ship for three months, in the middle of a leadership reshuffle. Here is what 3.7 Flash actually delivers, what it costs after the intro window, and why the model Google keeps shipping is not the one that matters most.

Read story →
alibaba qwen

Alibaba says Qwen 3.8 Max beats GPT-5.6 Sol. Independent evals put it 10th — and you may not be licensed to use it.

Qwen 3.8 Max shipped 3 August: 2.4 trillion parameters, 1M context, $2/$6 per million tokens. Alibaba's own benchmarks show it beating GPT-5.6 Sol and Claude Opus 4.8. Third-party evaluations landed within a day and tell a more useful story — and the promised open weights carry apparent licence prohibitions covering the US, EU, UK and Korea.

Read story →
deepseek open-weights

DeepSeek's cheap tier just beat its own flagship — and a pricing change is coming that hits Europe hardest

DeepSeek-V4-Flash-0731 shipped 31 July with the same architecture as the April preview and gains entirely from re-post-training. It outscores V4-Pro-Preview on all seven reported benchmarks at roughly a third of the price. The benchmarks can't be independently reproduced, and a peak-hours pricing policy is coming that doubles cost during European working hours.

Read story →
github copilot

GitHub Copilot moves to token-based billing — flat-rate is dead, power users seeing 10-50× bill spikes

GitHub Copilot transitioned all plans to usage-based billing on June 1, 2026. Base subscription prices unchanged (Pro $10, Pro+ $39, Business $19/user, Enterprise $39/user), but the monthly fee now buys an equivalent dollar amount of AI Credits ($10 plan = $10 in credits). Tokens consumed by every input, output, and cached request count against the allotment. Code completions and Next Edit stay free. Power users running agentic Copilot workflows report 10-50× bill increases vs the old flat-rate model.

Read story →
pricing deepseek

DeepSeek slashes V4-Pro prices 75% — now permanent (was promo through May 5)

Eight days after V4 Pro shipped, DeepSeek announced a 75% promotional cut on V4-Pro: input drops from $1.74 to $0.435 per million tokens, output from $3.48 to $0.87. Cache-hit input charges fall 90% across the entire DeepSeek API. Update May 22-25, 2026: the 75% cut is now PERMANENT — DeepSeek confirmed the discounted rates will not roll back after the originally-planned May 31 expiry.

Read story →
coding pricing

The $20 AI coding tier is quietly collapsing

This week GitHub paused Copilot Pro signups, Anthropic briefly removed Claude Code from the $20 Pro plan, and OpenAI consolidated Codex around higher tiers. The cheap agentic coding era is ending. Here's what it means for developers.

Read story →