AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Sep 3, 2026
·
geminipricingcodingagentsstrategy

Gemini 3.8 Flash is Google's fourth Flash model in under four months — and every one of them doubles in price on the same day

TL;DR: Google shipped Gemini 3.8 Flash on 2 September 2026 — its fourth Flash model in under four months. Introductory pricing is $0.75/$3.75 per million input/output tokens, with cached input discounted 90%, a 1M-token context window and 64K text output. Capability moves modestly: Artificial Analysis Intelligence Index 59 on high reasoning, up 3 points from 3.7 Flash. The finding: those introductory numbers are identical to what 3.7 Flash launched with on 13 August — and so is the expiry. Both revert to $1.50/$7.50 on 1 January 2027, a 100% increase. The discount is attached to the calendar, not the model, so upgrading buys you no extra runway. A restricted Flash Cyber variant also shipped, gated behind Google’s Fairwind programme. For you: budget the doubling now, and stop upgrading Flash on Google’s cadence.

What shipped

Gemini 3.8 Flash arrived on 2 September 2026. It accepts text, images, audio and video, carries a 1M-token context window with 64K-token text output, and Google reports improvements over 3.7 Flash in software engineering, agentic tasks and multi-step reasoning, with 54.9% on HLE-Verified.

Independent measurement from Artificial Analysis puts it at an Intelligence Index of 59 on high reasoning — up three points from 3.7 Flash — with 57 at medium and 52 at low. Throughput is roughly 300 output tokens per second on high reasoning. Cost per task lands at $0.58 high, $0.41 medium, $0.24 low, and time per task runs 2.5 minutes high to 0.8 minutes low.

It is live for Google AI Pro and Ultra subscribers in the Gemini app, in AI Mode in Google Search and in Gemini for Google Sheets, and for developers through Google Antigravity, AI Studio, the Gemini API and Android Studio.

That is a solid, unremarkable release. The interesting part is the price tag, and specifically what happens when you put it next to the last one.

The same price, and the same deadline

Gemini 3.8 Flash launched at $0.75 per million input tokens and $3.75 per million output, described as introductory and expiring 31 December 2026. From 1 January 2027 the rates become $1.50/$7.50.

Gemini 3.7 Flash launched on 13 August at $0.75/$3.75, described as a 50% introductory cut, expiring 31 December 2026, reverting to $1.50/$7.50.

Those are not similar numbers. They are the same numbers, with the same expiry date, three weeks apart.

The implication is worth stating plainly, because it inverts how introductory pricing normally works. A launch discount usually runs for a fixed window from the launch — adopt later, get the discount later. Here the window is anchored to a calendar date that both models share. Gemini 3.7 Flash buyers got roughly 4.5 months of discounted pricing. Gemini 3.8 Flash buyers get roughly four. The newer model comes with less runway, not more.

So the migration argument that usually accompanies a new release — move to the current version and reset your pricing clock — does not exist here. There is no clock to reset. Whichever Flash version your team standardised on, the bill doubles on 1 January 2027, and adopting 3.8 Flash today does not move that date by a single day.

Four models, one quarter

This is Google’s fourth Flash release in under four months: 3.5 Flash-Lite, then 3.6 Flash with its cyber variant in July, then 3.7 Flash in August, now 3.8 Flash. Each has been a genuine if incremental improvement, and each has cost adopting teams a re-qualification cycle.

Meanwhile the model Google actually promised at I/O in May — the Gemini 3.5 Pro flagship — has now missed its third deadline, with the base model scrapped. The pattern we flagged in August holds: Google keeps shipping the tier it can ship quickly, and the tier it committed to keeps not arriving.

For buyers this produces a specific and underappreciated cost. Version churn at three-week intervals is not free even when every release is an upgrade. Prompt suites drift, evaluation baselines reset, agent scaffolding needs re-testing, and cached-context assumptions change. If you chase every Flash release you will spend more engineering time migrating than the three-point Intelligence Index delta returns.

The defensible policy is to upgrade on your cadence rather than Google’s — quarterly is reasonable — and to break that rule only when a release fixes something you have actually hit in production.

Where the money actually is

Before modelling the January increase, pull your token split. Cached input carries a 90% discount, and for long-context agents that repeatedly read the same codebase or document set, cache reads are where most of the bill lives.

This is the same structural point that ran through Anthropic’s Fable 5.1 cache-read cut this week: frontier pricing is no longer one number you compare across a row. It is a structure of separate rates, and a headline percentage is a claim about the shape of your traffic. A workload that is 80% cache reads will feel the doubling very differently from an output-heavy one, and those two teams should reach different conclusions about whether to move providers.

The broader context is that the floor keeps moving. The open-weight price floor dropped again in late August, DeepSeek’s Flash line continues to undercut on multimodal, and the price war’s direction flipped in August as server costs rose on memory prices. A 2027 doubling is not obviously out of line with where costs are heading — which is precisely why treating it as a bluff would be a mistake.

The Cyber variant

Alongside the standard release, Google shipped Gemini 3.8 Flash Cyber — the same foundation model with more permissive cyber mitigations, tuned toward defensive work. Google reports 86.2% on CyberGym vulnerability discovery and 47.2% pass@1 on CWE-Bench patching, against Claude Fable 5’s 47.8%, and says it prioritised vulnerability fixing over offensive capability.

Access is application-gated through a programme called Fairwind, open to trusted government authorities, critical-infrastructure operators and software maintainers. Pricing for approved users is identical to the standard model.

That makes three restricted tiers announced by three labs inside a week, which we cover separately in the frontier-capability-tiers piece. The short version: the capability you can access is increasingly a question of who you are rather than what you pay.

What to do

Budget the doubling as your base case. Take your current Gemini Flash spend, double the input and output components, leave cached input at its discounted rate, and put that number in your 2027 plan. Treat any extension of the introductory window as upside rather than expectation.

Do not upgrade because a new Flash exists. There is no pricing benefit to moving from 3.7 to 3.8 — the rates and the expiry are identical. Move when a capability gap bites, and batch your migrations.

Instrument your token split now. You cannot evaluate the January increase, or any competing offer, without knowing your fresh-input, cached-input and output proportions. This is a one-afternoon task that will inform every pricing decision you make for the next two quarters.

Keep the exit open. Gemini and Gemini CLI are strong at this price point, and our AI coding tools shortlist covers the alternatives if the 2027 rate changes your calculus. A neutral gateway keeps that switch cheap.

The bottom line

Gemini 3.8 Flash is a good model at a good price, and the price has a date on it that most teams have not written down. The detail that matters is not the three-point capability gain — it is that Google has now launched two consecutive Flash models at identical introductory rates expiring on the same day, which tells you the discount belongs to the calendar rather than to any particular release.

Plan for $1.50/$7.50 from 1 January 2027. If Google extends the window, you will be pleasantly wrong. If you plan for the discount to renew and it does not, you will discover a 100% cost increase in the same week everyone else does.

Update, 3 September 2026 — a cheaper rate appeared the following day, with a different kind of price attached. Meta’s Muse Spark 1.3, released 2 September, lists a contributor endpoint at $0.10/$0.20 per million tokens — well under Gemini 3.8 Flash’s $0.75/$3.75 introductory rate, and under the $1.50/$7.50 this article tells you to budget for from 1 January. It is not a like-for-like substitute for two reasons. The contributor endpoint permits Meta to use your prompts and outputs to improve its products, which Flash’s standard terms do not; and Muse Spark’s own private-data tier is $1.25/$4.25, which is the honest comparison against Flash and lands close to the post-January Gemini rate rather than below it. The planning advice above is unaffected: budget for $1.50/$7.50 in January. But if the January cliff is what pushes you to shop around, note that the cheapest visible number in the market this week is cheap because it is partly paid in data. The arithmetic on Meta’s two tiers.

Frequently asked questions

Should I upgrade from 3.7 Flash to 3.8 Flash?

On price, it is a free move — the two carry identical rates and identical expiry, so there is no cost argument in either direction. On capability, the gain is real but modest: Artificial Analysis puts 3.8 Flash at 59 on its Intelligence Index for high reasoning against 56 for 3.7 Flash, a three-point move, with gains concentrated in software engineering, agentic tasks and multi-step reasoning. The reason to be deliberate rather than automatic is evaluation cost. This is the fourth Flash model since roughly June, and if you re-qualify your prompt suite, your evals and your agent scaffolding every three weeks you will spend more engineering time on migration than the capability delta returns. A reasonable policy is to upgrade Flash versions on a fixed cadence — quarterly, say — rather than on Google's release cadence, and to make an exception only when a release fixes something you have actually hit.

What exactly happens on 1 January 2027?

Input goes from $0.75 to $1.50 per million tokens and output from $3.75 to $7.50 per million — a 100% increase on both, applying to Gemini 3.7 Flash and 3.8 Flash alike. Google has published these as the standard rates all along; the launch prices are explicitly introductory. What catches teams out is the assumption that a discount which arrived with a new model will renew with the next one. There is no evidence for that yet, and the fact that 3.7 and 3.8 share an expiry date rather than each getting their own window is the clearest signal that the discount is attached to the calendar, not to the release. Budget for the doubling as the base case and treat any extension as upside.

Does the 90% cached-input discount change the maths?

Substantially, if your traffic has the right shape, and not at all if it does not. Cached input tokens receive a 90% discount, which for a long-context agent that repeatedly reads the same large codebase or document set is where most of the saving lives. That is the same structural point we made about Anthropic's cache-read cut: the headline rate is one line item among several, and vendors increasingly discount the line that suits their infrastructure. Before you model the January increase, pull your actual split of fresh input, cached input and output for the last month. A workload that is 80% cache reads experiences the doubling very differently from one that is output-heavy, and the two teams should reach different conclusions about whether to move.

Is Gemini 3.8 Flash good enough to replace a frontier model for coding?

For a lot of routine agentic and coding work, yes — that has been the Flash line's argument since 3.6 and it has only strengthened. It takes text, images, audio and video, has a 1M-token context window with 64K text output, and Google reports improvements in software engineering and multi-step reasoning, with 54.9% on HLE-Verified. Artificial Analysis measures roughly 300 output tokens per second on high reasoning and a cost per task of $0.58 high, $0.41 medium, $0.24 low. What it does not replace is the top of the range for work where a wrong answer is expensive to catch. The honest framing is tiering rather than replacement: route bulk and mechanical work to Flash, keep review, architecture and anything customer-facing on a frontier model, and measure the split rather than guessing at it.

What is Gemini 3.8 Flash Cyber and can I use it?

Almost certainly not, unless you are a government body, a critical-infrastructure operator or a software maintainer. Flash Cyber shares the same foundation model as standard 3.8 Flash but ships with more permissive cyber mitigations, tuned toward defence — Google says it prioritised vulnerability fixing over offensive capability, and reports 86.2% on CyberGym vulnerability discovery and 47.2% pass@1 on CWE-Bench patching, roughly level with Claude Fable 5's 47.8%. Access runs through an application-gated programme called Fairwind and is not publicly available. Pricing, if you are approved, is identical to the standard model. This is the third such tier to appear in a week across the three US frontier labs, which is a story in its own right.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.