AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.

Migrating Off Gemini 2.5: What to Move To, and What It Costs

Updated: Oct 6, 2026
7 tools · 2026

Gemini 2.5 Pro, Flash and Flash-Lite have no shutdown date and are not deprecated — but since 18 September 2026 Google only serves them to projects that already used them, so a new API key cannot call them at all. You are migrating for access, not for a deadline. Here is the replacement map, the real cost of each move priced on a reference workload, and the three gotchas Google's pricing page does not do the arithmetic on.

Gemini 2.5 is not being shut down, and that is exactly why you have to move. Google’s deprecation page lists gemini-2.5-pro, gemini-2.5-flash and gemini-2.5-flash-lite with no shutdown date and says they “are not deprecated.” But since 18 September 2026 Google is “limiting access to the 2.5 models to users who have actively used them in the past” — so an existing project runs indefinitely while a new API key or new Cloud project cannot call them at all. Google names two replacements: 3.5 Flash-Lite and 3.8 Flash. Neither is a like-for-like price. Our picks: 2.5 Flash → 3.8 Flash (budget 1.9x now, 3.9x from January); 2.5 Flash-Lite → 3.1 Flash-Lite, which is 30% cheaper than Google’s own recommendation but retires 7 May 2027; 2.5 Pro → nothing generally available.

Why you are migrating, and why no tracker warned you

The thing to understand before you plan anything is that this is an availability problem wearing a deprecation problem’s clothes.

A normal end-of-life gives you a date. You put it in a calendar, your dependency scanner flags it, procurement asks about it, and you migrate on your own schedule. Gemini 2.5 gives you none of that. The models sit on Google’s deprecation page with the shutdown column empty and a sentence saying they are not deprecated. Every automated check you own will tell you that you are fine.

What changed instead is who gets served. The notice on the 2.5 model pages reads, verbatim:

“To ensure reliable performance for everyone, we are limiting access to the 2.5 models to users who have actively used them in the past.”

It continues that the models “will continue to be served until further notice through the API,” and then: “For any new projects, use our latest models: 3.5 Flash-Lite or 3.8 Flash.”

That combination produces a failure mode worth naming, because it is the one that actually bites. Your production project keeps working. Your CI keeps passing. Your model string is still valid. Then someone provisions a fresh key — a new environment, a new customer tenant, a staging rebuild, a disaster-recovery restore — and that call fails against a model your dashboards insist is healthy. We covered the mechanism in detail when OpenAI retired the completions endpoint the same week; the fix is a monthly canary that provisions a genuinely new key and calls every model you depend on.

Two further details should kill any temptation to wait this out.

First, Google’s own page qualifies the dates it does publish: “The shutdown dates listed in the table indicate the earliest possible dates on which a model might be retired.” A published date is a floor, not a promise — which cuts both ways, but means you cannot treat an empty cell as safety.

Second, one 2.5 model does carry a real date. gemini-2.5-flash-image is listed for shutdown on 15 March 2027, with gemini-3.1-flash-lite-image as the named replacement — the only member of the 2.5 family with a published retirement date at all. Read that as the shape of what is coming for the rest: the family is not immortal, the text models are simply still undated.

The replacement map

You are onGoogle recommendsCheaper optionWhat to know
2.5 Flash-Lite ($0.10/$0.40)3.5 Flash-Lite ($0.30/$2.50)3.1 Flash-Lite ($0.25/$1.50)The cheap one is the one with a shutdown date: 7 May 2027. Thinking goes from Off to always-on.
2.5 Flash ($0.30/$2.50)3.8 Flash ($0.75/$3.75)3.6 or 3.7 Flash (same rates)All three Flash versions share the 1 Jan 2027 doubling to $1.50/$7.50. Picking an older one buys no runway.
2.5 Pro ($1.25/$10.00)not named3.1 Pro Preview ($2.00/$12.00)No GA successor exists. Gemini 3.5 Pro was scrapped and restarted in July 2026.
2.5 Flash-Image3.1 Flash-Image Preview—Already gone — shut down 2 October 2026.
2.0 Flash / Flash-Lite——Shut down 1 June 2026. If you are still on these, you are not calling Gemini.

Prices are per million tokens, input/output. Prices verified 4 October 2026 against the Gemini API pricing page and the deprecations page.

The row that deserves a second look is the first one. Google’s recommended Flash-Lite costs 42% more on output than the Flash-Lite it supersedes — $2.50 against 3.1 Flash-Lite’s $1.50 — and the model it supersedes is the one with a published retirement. That is an unusual shape, and it means the right answer depends on how expensive a model swap is for you rather than on which model is newer.

What the move actually costs

Rate cards do not tell you what you will pay, so here is a fixed workload to price against: 50,000 requests a month at 4,000 input and 600 output tokens each — 200M input, 30M output. That is the shape of a classification, extraction or support-triage service, which is what most 2.5 Flash-Lite traffic actually is.

ModelInput costOutput costMonthlyvs. your 2.5 tier
2.5 Flash-Lite$20.00$12.00$32.00baseline
3.1 Flash-Lite$50.00$45.00$95.003.0x
3.5 Flash-Lite (recommended)$60.00$75.00$135.004.2x
2.5 Flash$60.00$75.00$135.00baseline
3.8 Flash (through 31 Dec 2026)$150.00$112.50$262.501.9x
3.8 Flash (from 1 Jan 2027)$300.00$225.00$525.003.9x
2.5 Pro$250.00$300.00$550.00baseline
3.1 Pro Preview$400.00$360.00$760.001.4x

Two observations fall out of that table that no vendor page will hand you.

Google’s newest Flash-Lite costs exactly what the old mid-tier Flash cost. 3.5 Flash-Lite and 2.5 Flash both land on $135 on this workload. If you were on 2.5 Flash-Lite because it was the cheap tier, the model Google points you at is not a cheap tier any more — it is priced where the tier above used to be.

The Flash doubling is already on the rate card, in Google’s words. The pricing page states 3.8 Flash input as “$0.75 through December 31, 2026. $1.50 starting January 1, 2027,” and output as “$3.75 through December 31, 2026. $7.50 starting January 1, 2027.” The same two numbers and the same expiry apply to 3.7 Flash and 3.6 Flash. As we found when 3.8 Flash shipped, the discount is attached to the calendar rather than to the release, so there is no version of Flash you can adopt to postpone it. Budget the doubling as your base case and treat an extension as upside.

The thinking-token trap

This is the line item that makes real bills exceed the table above, and it hits the Flash-Lite migration hardest.

Gemini prices output as “output price (including thinking tokens)” — reasoning bills at the output rate. The default thinking behaviour changes across the migration:

So the cheapest-tier migration moves you from a model that emitted only the answer to one that emits reasoning you pay for at output rates and cannot disable. The 4.2x in the table assumes output stays at 600 tokens. It will not.

Sized on the reference workload, every 500 thinking tokens per request adds 25M output tokens a month:

Added per request3.5 Flash-Lite3.8 Flash (2026)3.8 Flash (2027)
+500 tokens+$62.50+$93.75+$187.50
+1,000 tokens+$125.00+$187.50+$375.00
+2,000 tokens+$250.00+$375.00+$750.00

At a modest 500 extra tokens, the 3.5 Flash-Lite move is $197.50 a month against a $32 baseline — 6.2x, not 4.2x. Do not estimate this. Run 200 real requests, read the thinking-token count off the usage metadata, and set thinking_level to the lowest value your evals tolerate rather than inheriting the default.

The grounding quota cut

If you use Grounding with Google Search, the migration changes your economics in a direction that depends entirely on your volume — and the direction is not the one the headline suggests.

1,500 a day is roughly 45,000 a month. So the shape is a cheaper meter bolted to a far smaller free bucket, and the crossover sits near 72,000 grounded requests a month. Below that the 2.5 arrangement was cheaper; above it, 3.x is.

The painful band is the middle. A team running 45,000 grounded requests a month paid $0 on 2.5 and pays about $560 a month on 3.x — a line item that appears from nothing, in a migration sold as a price cut. A team running 150,000 saves roughly $830 a month. Check which side you are on before you forecast.

Cache pricing moves similarly: cache storage on 3.8 Flash improves to $0.50 per million tokens per hour from 2.5 Flash’s $1.00, while cache reads rise from $0.03 to $0.075, and to $0.15 in January. A cache-heavy agent and a cache-light classifier should reach different conclusions.

If you are on 2.5 Pro, there is no like-for-like move

This is the case with no clean answer, and it is worth being blunt about it.

Google has shipped four Flash models since roughly June 2026 — 3.6, 3.7, 3.8 and the Flash-Lite variants — and in the same window has shipped no generally available Pro model. The newest Pro on the rate card is Gemini 3.1 Pro Preview. Gemini 3.5 Pro missed three launch targets before Google scrapped a nearly-finished base model and restarted pretraining.

So a 2.5 Pro user is asked to move from a GA model to a preview model, at 1.4x the cost, with no GA successor dated. Three workable responses:

  1. Stay, if you can. Your existing project still serves 2.5 Pro. The access gate blocks new projects, not yours. This is the one case where waiting is defensible — but pair it with the new-key canary, because you are now depending on an undated arrangement.
  2. Move the workload down, not across. Much 2.5 Pro traffic does not need a Pro model. 3.8 Flash at an Artificial Analysis Intelligence Index of 59 on high reasoning is a serious model, and on this workload it is half the price of 2.5 Pro even after January.
  3. Qualify a second vendor. If you genuinely need frontier reasoning, the honest read is that Google does not currently sell a GA one. Claude and ChatGPT do.

Note that the pressure is not only on the API. Google is also cutting the Gemini app’s free tier to Flash-Lite and removing the Pro model from the $4.99 AI Plus plan from 9 October — which removes the easiest way to evaluate a Gemini model before committing code to it.

Migration checklist

  1. Inventory the strings, not the services. Grep for gemini-2.5 across application code, notebooks, cron jobs, Terraform, CI config and vendor SDK defaults. Legacy model strings outlive the code that was supposed to own them.
  2. Provision a new key and test it now. This is the single highest-value step, because it tells you whether you are already exposed. A new key in a new project calling gemini-2.5-flash is the exact failure your monitoring cannot see.
  3. Measure thinking tokens before you price anything. 200 real requests, read the usage metadata, then use the sensitivity table above.
  4. Pull your grounding volume. Find out which side of ~72,000 requests a month you are on.
  5. Split the Flash-Lite decision on migration cost. Cheap to change a model string later → 3.1 Flash-Lite and re-pick before 7 May 2027. Expensive → 3.5 Flash-Lite now.
  6. Re-run evals, not just the bill. Thinking defaults change latency as well as cost; a minimal-thinking Flash-Lite is a different latency profile from a thinking-off one, which matters for anything user-facing.
  7. Make the canary permanent. Monthly, provision a genuinely new key and call every model your stack depends on. Alert on any 404. It is a twenty-line job and it is the only check that catches a soft access gate.

For the broader picture on which models are worth standardising on, see our best AI coding tools and the Gemini review; for the terminal side of the Google stack, Gemini CLI. High-volume teams pricing an exit should also read DeepSeek and Qwen.

What changed

Frequently asked questions

When is Gemini 2.5 being shut down?

It is not, and that is the confusing part. Google's deprecation page lists gemini-2.5-pro and gemini-2.5-flash (both released 17 June 2025) and gemini-2.5-flash-lite (22 July 2025) with no shutdown date announced, and states plainly that these models 'are not deprecated and will continue to be served until further notice through the API.' Third-party deprecation trackers circulating 16 or 20 October 2026 dates for gemini-2.5-pro are not quoting Google. What exists instead is an access gate: since 18 September 2026 Google is 'limiting access to the 2.5 models to users who have actively used them in the past.' Existing projects keep working indefinitely; new ones cannot start. Note also that the shutdown dates Google does publish are described on its own page as 'the earliest possible dates on which a model might be retired' — so even a listed date is a floor, not a commitment.

What does Google say to migrate to?

Two models, named explicitly on the deprecation page: 'For any new projects, use our latest models: 3.5 Flash-Lite or 3.8 Flash.' Neither is a like-for-like price replacement. On a 200M-input / 30M-output monthly workload, 2.5 Flash-Lite costs $32 and the recommended 3.5 Flash-Lite costs $135 — 4.2x — while 2.5 Flash at $135 becomes $262.50 on 3.8 Flash today and $525 on 1 January 2027 when every 3.6/3.7/3.8 Flash rate doubles. If you are on 2.5 Pro there is no generally available successor at all: the newest Pro model on the rate card is Gemini 3.1 Pro Preview, and Gemini 3.5 Pro was scrapped and restarted in July 2026.

Is Gemini 3.5 Flash-Lite really the cheapest Flash-Lite?

No — Gemini 3.1 Flash-Lite is cheaper on both halves of the rate card, at $0.25 input and $1.50 output per million tokens against 3.5 Flash-Lite's $0.30 and $2.50. On the reference workload that is $95 a month against $135, so Google's recommended model costs 42% more than the one it supersedes. The catch is the reverse of the usual one: 3.1 Flash-Lite is the version with a published shutdown date, 7 May 2027, while 3.5 Flash-Lite has none. So the cheap option is the one that expires and the expensive option is the one that does not. If your workload is price-sensitive and your model string is a single config value, 3.1 Flash-Lite buys roughly eighteen months at a 30% discount and you re-pick in early 2027. If migration is expensive for you, pay the 42% and go straight to 3.5 Flash-Lite.

Why did my bill go up more than the rate card predicted?

Almost certainly thinking tokens. Gemini bills 'output price (including thinking tokens)' on every model, and the default thinking setting changed underneath you. Gemini 2.5 Flash-Lite is documented with thinking Off by default; Gemini 3.5 Flash-Lite ships On at minimal and its supported levels are minimal, low, medium and high — there is no off. Gemini 3.8 Flash defaults to medium. So a model that previously emitted only the answer now emits reasoning you pay output rates for and cannot switch off. On the reference workload, every 500 thinking tokens added per request is 25M extra output tokens a month: +$62.50 on 3.5 Flash-Lite, +$93.75 on 3.8 Flash this year, +$187.50 on 3.8 Flash from January. Measure your actual thinking-token count on a sample before you size the budget, and set thinking_level to the floor your evals tolerate rather than inheriting the default.

Does anything get cheaper by moving to 3.x?

Two things, and one of them is a trap. Grounding with Google Search drops from $35 to $14 per 1,000 requests, a 60% cut — but the free allowance goes from 1,500 requests per day on 2.5 to 5,000 per month shared across all 3.x models, which is roughly an 89% reduction in free volume. The crossover is around 72,000 grounded requests a month: below it the 2.5 arrangement was cheaper, above it 3.x is. A team running 45,000 grounded requests a month paid nothing on 2.5 and pays about $560 a month on 3.x. The other genuine improvement is cache storage on 3.8 Flash at $0.50 per million tokens per hour against 2.5 Flash's $1.00, though cache reads themselves rise from $0.03 to $0.075 and then $0.15 in January.

Should I move to a different vendor instead?

It is worth pricing, because the 2.5 gate is the second time in a month that a lab has removed a cheap tier without a deprecation entry, and the January Flash doubling is already published. The specific case for looking elsewhere is the Flash-Lite tier, where the move Google recommends is a 4.2x increase: at that point DeepSeek's and Qwen's high-volume rates are in the same conversation rather than obviously worse, and open-weight models you host yourself start to pencil out on steady traffic. The case against is that nobody matches Gemini's 1M-token context and Google Search grounding as an integrated package, and a vendor switch costs far more engineering time than a model-string change. Price the alternative, but do not assume the migration is free on either path.