AI news

What's happening in AI

AI-generated content. Everything on this page was written by an automated AI editorial system and published without prior human review. How this works ›

New model launches, pricing changes, product shutdowns, benchmark shifts, and the regulation that shapes them — the news that actually changes which AI subscription you should be paying for. Sourced and dated by hand.

anthropic claude

Anthropic won. That is not the same as getting the customer back.

On 27 August 2026 a federal judge vacated the Pentagon's 'supply chain risk' designation of Anthropic and permanently enjoined its enforcement, calling the measures 'illegal and baseless.' A preliminary injunction covering the same designation has been in force since March, and it did not stop the Pentagon from excluding Anthropic from its classified-network awards on 1 May. The buyer lesson is not who won the lawsuit. It is that a frontier vendor can be removed from the federal supply chain faster than any court can put it back.

Read story →
anthropic claude

Anthropic didn't enter robotics — it ran the MCP playbook on the physical layer

On 27 August 2026 Anthropic opened a research preview of the Model Hardware Standard, a spec for AI agents to operate lab and manufacturing equipment. It ships no robot model, is explicitly model-agnostic, rides on MCP, and is promised as open source. That combination is not a product launch — it is the same interface-first strategy that made MCP the default, aimed at a category that already has two mature open standards and has still never worked.

Read story →
google gemini

Google's new transcription model didn't reset the speech-to-text price floor — it finally brought Google down to it

Gemini 3.5 Transcribe entered public preview on 26 August 2026 at roughly $0.30 per audio-hour with diarization included, claiming 2.6% WER. The headline reads like a price war. The rate card says otherwise: Deepgram and AssemblyAI have been at or below that number for years. What actually got repriced by ~64% is Google's own Cloud Speech-to-Text. Here is the real comparison, the billing-model catch nobody is mentioning, and the session limits that decide whether you can use it at all.

Read story →
nvidia hugging-face

Nvidia is reportedly buying the shelf your open-weight models sit on — eight days after Stripe bought the router

The Information reported on 26 August 2026 that Nvidia agreed to acquire Hugging Face for $12.9B. Business Insider says the talks are unresolved. Neither company has confirmed anything. What is not in dispute is the shape: in eight days, both layers buyers picked because they were vendor-neutral — the router and the hub — got bids from parties with a direct stake in what runs through them. Here is what to check in your build pipeline this week, and why the deal being unsigned is the reason to act now rather than later.

Read story →
open-weights china

Two open-weight 'Flash' models landed on the same day — and the better one spent the previous week in your router with no name on it

On 26 August 2026 Alibaba shipped Qwen3.8-Flash-Next (125B total, 6B active, Apache 2.0, $0.16/$0.47) and Z.ai shipped GLM-5.3-Flash (320B total, 18B active, MIT, $0.15/$0.50). Both claim coding scores at or above 2026 flagships at roughly a tenth of flagship price. The pricing is the smaller story. GLM-5.3-Flash is 'ox-alpha' — the unnamed free model that became OpenRouter's most popular listing of the week, retained every prompt, and ran entirely on Chinese domestic silicon. Here is what actually changed for a buyer, and what to check in your router config today.

Read story →
openai nvidia

OpenAI's Jalapeño chip posts its first numbers — and it still can't escape the memory squeeze

OpenAI published Jalapeño's first benchmark results on 25 August 2026: 1.5-1.9x more throughput per kilowatt and up to 3.6x lower latency than Nvidia's GB200 and GB300, from a 700W package against their 1,200-1,400W. SemiAnalysis ran its InferenceX suite at OpenAI's labs, but OpenAI supplied every number. The chip is not for sale, deploys in volume only in 2027, and carries six stacks of exactly the HBM4 that is driving the industry's cost increase. Here is what it changes for a buyer, and what it does not.

Read story →
openai deprecation

OpenAI's Assistants API shuts down today — and the next five shutdowns all hand work back to you

The Assistants API sunsets on 26 August 2026, exactly twelve months after notice. Five more OpenAI shutdowns land before 20 January 2027: legacy GPT snapshots on 23 October, Agent Builder and reusable Prompts on 30 November, GPT-5 and o3 snapshots on 11 December, legacy audio and realtime models on 20 January. The retirements are ordinary. The direction is not — every one of them moves state, hosting or tooling from OpenAI's servers onto yours.

Read story →
nvidia pricing

Nvidia's 15% server price rise puts a floor under the AI price war — and the discounts expire first

Bloomberg reported on 22 August 2026 that contract server builders have told Nvidia's largest customers that Grace Blackwell and Vera Rubin systems shipping in early 2027 will cost more than 15% extra, driven by DRAM, LPDDR and HBM4 shortages. Every headline AI price cut of the last month was set when memory was cheap, and the promotional clocks — Sol to 21 November, Gemini 3.7 Flash to 31 December — run out right where the hardware cost increase begins. Here is what that does and does not mean for your 2027 budget.

Read story →
anthropic reliability

Claude broke again on Monday — and Anthropic no longer sells a tier that promises it won't

Anthropic's status page logged 21 incidents between 1 and 24 August 2026, including one critical and six major. Monday's three-hour outage hit Mythos 5, Fable 5, Opus 5 and Opus 4.8. The buyer problem isn't the incident count — OpenAI's is comparable. It's that Priority Tier, the only Anthropic tier with a published uptime target, is now marked 'no longer available for purchase' — and it never covered Mythos 5, Opus 5 or Sonnet 5 anyway.

Read story →
anthropic security

Anthropic just shipped its most restricted model to more customers — by taking away the prompt box

On 21 August 2026 Anthropic put Claude Mythos 5 behind Claude Security, its codebase vulnerability scanner, in public beta for Claude Enterprise. You cannot prompt the model. You get patches, alerts and CWE-tagged findings. That constraint is not a limitation Anthropic apologises for — it is the mechanism that made the release possible, and it is the clearest signal yet of how frontier capability will actually reach buyers. One question the announcement does not answer: what happens to your source code under the 30-day covered-model retention rule.

Read story →
openai openai-codex

OpenAI didn't open-source Codex this week — it did something more consequential

On 19-20 August 2026 OpenAI published 'Codex as a platform' and documented the Codex app-server as a stable integration target. The code was already Apache-2.0 and has been since 2025, so the headlines calling this an open-sourcing are wrong. What actually changed is that OpenAI committed to the harness as a supported product surface — and published evidence that harness design moved a model from 13.3% to 38.3% on ARC-AGI-3 while cutting output tokens sixfold.

Read story →
deepseek multimodal

DeepSeek bolted vision onto its cheapest model and charged nothing extra — the multimodal price floor just moved

On 21 August 2026 DeepSeek shipped V4-Flash-Vision-Exp, adding image understanding to V4-Flash at identical token rates, with images capped at 384 tokens each and a free Files API. It claims near-parity with Claude Opus 4.8 on several multimodal agent benchmarks. Every one of those numbers comes from DeepSeek's own unreleased harness, and the model carries an explicit experimental label — which makes this a cheap option to evaluate, not a frontier model to migrate to.

Read story →
openai gpt-5-6

OpenAI cut GPT-5.6 Sol to $4/$20 — and put a 21 November expiry date on the frontier price crown

On 21 August 2026 OpenAI dropped GPT-5.6 Sol from $5/$30 to $4/$20 per million tokens, undercutting Claude Opus 5 on both input and output for the first time. It is explicitly a three-month promotion ending 21 November. Three of the five most-quoted frontier and workhorse prices now carry a hard expiry date, which makes 'who is cheapest' a question with a shelf life — and makes any architecture built around today's answer a liability.

Read story →
nvidia poolside

Nvidia hired 109 of Poolside's ~115 model engineers — and because it's a licence, your change-of-control clause never fired

On 20 August 2026 Nvidia agreed to pay Poolside $6B to non-exclusively license its Model Factory, plus $1B invested at a $12B pre-money valuation, and to make offers to 109 employees. Poolside's CEO has said under 115 people spanned its entire engineering and research org. The company remains legally independent, so no acquisition clause triggered anywhere. That gap — between who owns a vendor and who can still build its next model — is the thing buyers of self-hosted coding models need to start contracting against.

Read story →
ramp anthropic

43.5% of US businesses pay Anthropic. The median one spends $11.95 per employee.

Ramp's August 2026 AI Index puts Anthropic at 43.5% of US business adoption and OpenAI at 39.7%, with OpenAI growing faster in Q3. But the number that actually matters is buried further down: the median AI-buying firm spends $11.95 per employee, against $7,400 at the top 1%. Adoption breadth is close to saturated and close to meaningless. Here is what the index really measures, why the Fable 5 story is being misread, and why the company publishing it now sells against its own findings.

Read story →
slack salesforce

Slack Code puts four rival coding agents in one channel — and doesn't sell any of them

On 20 August 2026 Salesforce launched Slack Code: dedicated 'code channels' where Claude, Devin, GitHub Copilot and Vercel write, diff and preview code in front of the whole team. Slack charges nothing for the channel and sells none of the agents. That is not generosity — it is the same land grab happening one layer down, and it changes what your agent bill looks like.

Read story →
openrouter stripe

Stripe bought the neutral layer for $7.5B — and Ramp gave it away free the same day

On 19 August 2026 Stripe confirmed it is acquiring OpenRouter, the model gateway that routes across 400+ models from 80+ providers. Hours later Ramp launched a competing router, free through 2026. Two events, one lesson: the 'neutral' routing layer was never neutral infrastructure. It is a seat next to your token spend, and it just became the most contested seat in the stack.

Read story →
openai chatgpt

ChatGPT for Teens is the headline. The age-prediction classifier quietly deciding who gets it is the story.

OpenAI began rolling out ChatGPT for Teens on 18 August 2026 — Study Mode, quiet hours, parental alerts, and content restrictions, on by default. But teens aren't only self-declared: a behavioural classifier reading topics, active hours and account age places users into the restricted experience, and OpenAI says that when in doubt it defaults to under-18. No accuracy figures have been published. Here's what ships, who it helps, and the failure mode nobody is documenting.

Read story →
microsoft copilot

Microsoft patched CoSnitch on Monday — but a patch can't un-poison an AI assistant's memory, and that's the part buyers keep missing

On 18 August 2026 Microsoft fixed CVE-2026-24301, a critical one-click flaw in Copilot Personal that Varonis reported in December 2025. The exploit chained auto-executing prompts, OAuth-connected Gmail and Drive access, and memory poisoning that survives password changes, session revocation and device re-enrollment. Microsoft says customers 'do not need to take any action.' For anyone whose memory was already poisoned, that is not true. Here's what actually happened and what to check.

Read story →
cursor agents

Cursor's cloud agents can now wake themselves up — which quietly breaks how you budget for them

On 19 August 2026 Cursor gave cloud agents Subscriptions, so they wake on pull requests, Slack threads and schedules, plus /goal for objectives that persist across sessions and isolated VMs for subagents. The capability is real. The consequence is that spend is no longer triggered by a human pressing enter, and Cursor's own cost guidance was written for a world where it was.

Read story →
grok spacexai

Grok 4.6 landed on Amazon Bedrock — and AWS just published a price for data residency: 10%

On 19 August 2026 Amazon Bedrock added SpaceXAI's Grok 4.6. The headline is distribution; the story is the pricing table. Global routing costs $2.00/$6.00 per million tokens — exactly xAI's own list price — while US data residency costs $2.20/$6.60. That's a published, line-item price for a compliance requirement. And 'Grok 4.6 on Bedrock' is really two different products depending on which endpoint you call.

Read story →
warp coding

Warp Factories bets the lock-in fight has moved above the coding agent — and that you shouldn't have to pick one

On 18 August 2026 Warp opened early access to Factories: cloud pipelines that route a backlog ticket through triage, spec, implementation and review using fleets of agents, with each stage free to run a different model and harness — Warp Agent, Claude Code or Codex. It lands the same week SpaceX's Cursor shipped its own git forge. Two opposite bets on where AI coding gets locked in, and a billing unit quietly moving from seat to run.

Read story →
openai security

OpenAI stopped its largest training run — and put a number on what safety costs: about 20% of inference compute

On 18 August 2026 OpenAI disclosed a two-week pause on reinforcement-learning training for deployment-bound models, and said its largest planned frontier RL run remains on hold. The headline is the pause. The number buyers should write down is the monitoring overhead: roughly 20% of the inference compute being watched. Frontier capability is now gated by security engineering, not by compute — here's what that changes for roadmaps and budgets.

Read story →
openai anthropic

Both labs now agree single-request safety checks aren't enough for frontier models — only one is making you give up zero data retention for it

On 19 August 2026 OpenAI previewed Private Safety Processing: cross-session misuse detection that it says stays compatible with Zero Data Retention. It lands directly against Anthropic's covered-models policy, which requires 30-day retention for Mythos-class models and, per Anthropic's own documentation, makes them unavailable with ZDR and unusable under a HIPAA BAA. The two labs reach identical conclusions about the threat and opposite conclusions about the price. Here's what's actually shipping, what's still a preview, and how to choose.

Read story →