Google's 'Frozen v2' would etch Gemini's architecture into silicon — a 6–10× efficiency bet that model design has stopped moving
TL;DR: Google is reportedly developing an AI inference chip codenamed “Frozen v2” that etches Gemini’s architecture — the blueprint, not the weights — directly into silicon. Engineers project 6 to 10× more tokens per watt than Google’s latest TPUs, which would be among the largest single-generation efficiency leaps in AI inference hardware. Deployment is targeted for 2028, complementing rather than replacing TPUs, and the project reportedly aims to ease internal compute shortages that have limited Google Cloud’s ability to serve enterprise customers. Alphabet stock moved on the report (CNBC). The catch: it only pays off if future Gemini models keep the same architecture — this is a wager that model design has stopped changing fast. Important: all of this is reporting, not a Google announcement. What this means for you: nothing to buy, but it’s a strong signal that the next competitive front is inference cost per watt — which is what keeps pushing your price-per-token down.
What’s reported
Per Tom’s Hardware, The Decoder, Quartz, and CNBC, Google is working on a server chip internally called Frozen v2:
- It would embed elements of Gemini’s architecture directly into hardware, reducing the computation and data movement required during inference.
- Engineers project 6–10× more AI tokens per unit of power versus Google’s latest TPUs.
- It’s targeted for 2028 deployment and is meant to complement, not replace, the TPU line.
- The motivation is reportedly internal compute scarcity — shortages that have constrained Google Cloud’s ability to serve some enterprise customers.
- Alphabet’s stock rose on the report.
Google has not announced the chip, confirmed the performance figures, or commented on the timeline. Everything above is sourced reporting.
The design choice that actually matters
Most coverage led with “Google is hardwiring Gemini into a chip,” which is close enough to be misleading. The important detail is what gets frozen.
Frozen v2 reportedly embeds the architecture — the structural blueprint of how layers, attention, and data flow are organised — not the weights, the trained parameters that encode what a model actually knows. That distinction is the difference between a clever bet and an obviously bad idea:
- Freezing weights would produce a chip that runs exactly one model. The moment Google trained a better Gemini, the silicon would be scrap.
- Freezing architecture lets Google keep training and deploying new Gemini models on the same hardware — as long as those models keep the same fundamental shape.
That’s why the efficiency gain is plausible. A general-purpose accelerator spends enormous power shuttling data between memory and compute units because it has to accommodate any model. If you know in advance the exact shape of the computation, you can lay the silicon out to match it and eliminate much of that movement. This is the same logic that made ASICs crush GPUs for bitcoin mining — specialise the hardware to a known, stable computation.
Which sets up the real wager.
Why this matters
1. It’s a bet that frontier architecture has stopped moving. This is the thesis worth sitting with. Committing silicon in 2026 for a 2028 deployment, built around today’s Gemini architecture, is a statement that Google believes the fundamental shape of frontier models won’t change much in that window. For the last several years that would have been a reckless assumption — architectures shifted constantly. If Google is right, the industry has entered a maturity phase where the gains come from scale, data, and efficiency rather than structural reinvention. If Google is wrong — a new attention mechanism, a different sparsity regime, something not yet published — it ends up with beautifully efficient silicon optimised for a shape nobody builds anymore.
2. Inference, not training, is where the money leaks. Training a frontier model is a huge one-time cost; serving it is a permanent one that scales with every user query. As AI products move from novelty to infrastructure, tokens-per-watt becomes the metric that determines margin. A 6–10× improvement there wouldn’t just save Google money — it would change what’s economically viable to offer for free, which is why this is ultimately a product story wearing a hardware costume.
3. The compute-shortage motive explains a lot of Google’s year. If Google Cloud genuinely can’t serve some enterprise customers for lack of capacity, that context sits uncomfortably close to the three missed Gemini 3.5 Pro deadlines and the broader execution questions around DeepMind this summer. Capacity constraints don’t cause a scrapped base model, but a company rationing compute is a company making harder trade-offs everywhere. Worth holding as context, not as causation.
4. Custom silicon is now the whole industry’s second front. OpenAI has Jalapeño with Broadcom; Anthropic has expanded compute partnerships with Google and Broadcom; Google has TPUs and now, reportedly, Frozen v2. Every frontier lab has concluded that renting general-purpose GPUs forever is a losing position. For buyers, the practical read is that the labs with their own silicon will have more room to cut prices without cutting margin — which is the mechanism behind cheap tiers like GPT-5.6’s Terra and Luna.
5. It reframes the open-weight cost argument. Open-weight models like Kimi K3 and DeepSeek compete largely on price. If closed labs achieve step-change inference efficiency on proprietary silicon, part of that price advantage erodes — the closed model gets cheaper to serve at the same time the open one requires you to rent someone else’s general-purpose hardware. Nothing about this is settled, but it’s the counter-move to watch.
What “tokens per watt” actually buys
The metric in the reporting is deliberately not “faster.” It’s tokens per unit of power — and that framing tells you where the real constraint now sits.
Frontier AI is no longer primarily bottlenecked by chip supply. It’s bottlenecked by electricity and the datacentre capacity to deliver it. Power contracts, grid connections, and cooling are multi-year commitments that can’t be accelerated by spending more; a lab can buy chips faster than it can energise buildings to run them. In that world, a 6–10× improvement in tokens per watt isn’t a speed upgrade — it’s 6–10× more product served from the same physical footprint you already have permission to power.
That’s why the reported motivation (easing internal compute shortages that limit Google Cloud’s enterprise serving) is coherent rather than corporate-speak. If capacity is the wall, efficiency is the only way through it that doesn’t require new substations.
For buyers, this reframes a lot of 2026’s odd behaviour — usage caps, rate limits, staged rollouts, enterprise waitlists. Those aren’t purely commercial choices; they’re rationing under a physical constraint. Any technology that materially raises tokens per watt loosens that rationing, which shows up on your side as higher limits and lower prices rather than as a spec you ever read.
What this means for you
- Change nothing today. This is unconfirmed 2028 hardware. Buy on models and prices that exist now — see the best AI chatbots guide.
- Expect continued downward price pressure on inference. The whole industry is optimising the same variable. Plan budgets assuming cost-per-token keeps falling, not rising.
- If you’re a Google Cloud enterprise customer: the reported capacity constraints are the more actionable detail than the chip. If you’ve hit quota or provisioning friction, that’s consistent with the reporting — factor it into vendor redundancy planning.
- Don’t read the stock move as validation. Alphabet shares rising on a report reflects investor sentiment about a leak, not confirmation that the chip works or ships.
The honest caveats
- Google hasn’t confirmed any of it. No announcement, no spec sheet, no comment. The codename, the 6–10× figure, and the 2028 date all come from reporting citing internal sources.
- “6 to 10× more tokens per watt” is an engineering projection, not a measurement. Projections on unbuilt silicon routinely compress once real workloads, memory bandwidth, and thermals intervene. Treat it as an aspiration.
- 2028 is a long way out. Chip programmes slip, get rescoped, or get cancelled. Google has strong custom-silicon credentials with TPUs, but a two-year-out roadmap is not a product.
- The flexibility risk is real and unhedgeable. The entire value proposition depends on Gemini’s architecture remaining stable. That’s a genuine, non-trivial bet — and one Google is better positioned to make than most, since it controls both the model and the chip.
- Complementing TPUs means narrower impact than headlines imply. This isn’t a TPU replacement; it’s a specialised inference companion. The efficiency gain applies to a slice of the workload, not to everything Google runs.
The grounded summary: if Frozen v2 is real and lands, the interesting part isn’t the speed — it’s what building it says. Etching an architecture into silicon two years ahead of deployment is only rational if you think the shape of frontier AI has finally stopped moving. That’s a much bigger claim than any benchmark, and the next two years will settle it.
Frequently asked questions
What is Google's Frozen v2 chip?
Per reporting, Frozen v2 is an AI inference server chip Google is developing that embeds parts of Gemini's model architecture directly into the hardware. The goal is to cut the computation and data movement needed to serve a model, with engineers projecting 6 to 10 times more AI tokens per unit of power than Google's latest TPUs. Deployment is reportedly targeted for 2028, complementing rather than replacing TPUs. Google has not publicly announced it.
Does it hardwire Gemini's weights into the chip?
No — and this distinction is the whole design. Frozen v2 reportedly embeds the model architecture (the structural blueprint: how layers, attention, and data flow are organised) rather than the trained weights (the tuned parameters). That means Google can still train and swap in new Gemini models, as long as those models keep the same underlying architecture. Freezing weights would have made the chip obsolete the moment a new model shipped.
Why would Google do this?
Two reasons. First, efficiency: general-purpose accelerators spend a lot of power moving data around, and hardwiring the architecture removes much of that overhead. Second, scarcity — the project is reportedly aimed at easing internal compute shortages that have limited Google Cloud's ability to serve some enterprise customers. Inference, not training, is where the sustained cost of running AI products sits.
What's the catch?
Flexibility. A chip with an architecture etched into it only helps if future models keep that architecture. If frontier model design shifts meaningfully before 2028 — a new attention mechanism, a different sparsity approach — Google could be left with highly efficient silicon optimised for a shape nobody builds anymore. It's a bet that architecture innovation has slowed enough to be safely frozen.
Should this change what AI tools I buy today?
No. This is reported 2028 hardware, not a product you can buy, and Google hasn't confirmed it. The second-order effect is what matters: if inference efficiency jumps this much across the industry, cost-per-token keeps falling, which is already reshaping pricing tiers. Make decisions on today's shipping models and prices, not on a chip that's two years out.
Sources
- Google reportedly developing 'Frozen v2' chip with Gemini's architecture etched into the silicon (Tom's Hardware)
- Alphabet stock pops on report it's developing a more efficient AI chip (CNBC)
- Google's 'Frozen v2' chip reportedly bakes Gemini's architecture directly into silicon for efficiency gains (The Decoder)
- Google developing Gemini-specific chip called Frozen v2 (Quartz)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.