OpenAI's oldest API surface runs out of models on Monday — and its own docs disagree about when your legacy fine-tunes die
TL;DR: On Monday 28 September 2026 OpenAI shuts down gpt-3.5-turbo-instruct, davinci-002, babbage-002 and gpt-3.5-turbo-1106 — announced 26 September 2025, so a full year’s notice. Three of those four are the complete list of models served by the legacy /v1/completions endpoint, which therefore does not get a successor model; it gets zero. The deprecation table maps all four onto gpt-5.6-terra, a reasoning model on the Responses/Chat APIs at $2.00 in / $12.00 out against instruct’s $1.50 / $2.00 — 6x the output rate, before reasoning tokens, which also bill as output. echo, suffix, best_of and raw base-model log-probabilities have no successor anywhere in OpenAI’s catalogue. And the fine-tuning section contradicts itself: the stated rule is that fine-tune inference lasts “until the base models are deprecated” (28 September), while ft-babbage-002 and ft-davinci-002 are listed for 23 October. Assume Monday. Meanwhile Google has gated Gemini 2.5 to prior users since 18 September with no shutdown date and no deprecation entry — availability risk that no deprecation tracker can see.
What actually happens on Monday
OpenAI’s deprecations page lists four model IDs with a shutdown date of 28 September 2026: gpt-3.5-turbo-instruct, davinci-002, babbage-002 and gpt-3.5-turbo-1106. Each carries the same recommended replacement, gpt-5.6-terra. The announcement date on all four is 26 September 2025 — which makes today, as this publishes, the anniversary of the notice and the second-to-last business day before it lands.
Twelve months of notice is a fair clock. It is longer than the Assistants API got in spirit if not in length, and much longer than Sora’s users got. This desk is not going to complain about the runway.
The problem is the shape of the row, not the length of the clock. A deprecation table with a “recommended replacement” column communicates one specific thing: change this string to that string. For gpt-3.5-turbo-1106, a chat-era snapshot, that is roughly accurate. For the other three it is not accurate at all, and the reason is visible only if you open a different page.
An endpoint with no models left
OpenAI’s API reference for Create completion — the legacy /v1/completions route, the original 2020-era text-in-text-out surface — lists the models it supports. There are exactly three: gpt-3.5-turbo-instruct, davinci-002 and babbage-002.
All three shut down on Monday.
So the legacy completions endpoint is not losing a model and gaining a replacement. It is losing every model it has. After 28 September, the route remains documented, carries no deprecation banner, and has nothing you are permitted to ask it for. Every model value its own reference page names is dead.
This distinction matters because of how migrations get planned. A model deprecation is a configuration change and gets assigned to whoever owns the config. An endpoint removal is an integration change and needs an engineer who understands the calling code. OpenAI filed this as the former. For most affected workloads it is the latter, and the four weeks between now and the next cliff on 23 October are not enough time to discover that by accident.
The four parameters with nowhere to go
The completions endpoint carries parameters that exist nowhere else in OpenAI’s API, and they are not decoration:
echo— returns the prompt along with the completion. Paired withlogprobs, this is the mechanism for scoring arbitrary text: you feed in a string, echo it back, and read the per-token log probabilities to get a likelihood or a perplexity.suffix— fill-in-the-middle insertion between a prefix and a suffix. Supported only ongpt-3.5-turbo-instruct.best_of— generate n completions server-side, return the one with the best log probability. A quality lever that costs you nothing in round trips.logprobs— up to five most-likely tokens per position, with probabilities.
And underneath all of them, the thing that actually disappears: davinci-002 and babbage-002 are base models. Not instruction-tuned, not chat-formatted, not RLHF’d into a helpful assistant — raw next-token predictors. They were the last base models OpenAI exposed commercially, and for three years they were the only way to get a calibrated likelihood for a piece of text out of an OpenAI model.
A substantial amount of work rests on that: academic evaluation pipelines, calibration and uncertainty research, classifier confidence scores, prompt-selection tooling, membership-inference and contamination studies, and any scoring system that ranks candidate strings by model likelihood rather than by asking a chat model to rate them.
gpt-5.6-terra cannot do any of it. It is a reasoning model on the Responses and Chat Completions APIs. It does not do raw continuation, it rejects echo, suffix and best_of, and there is no parameter that makes it hand back a clean likelihood for a string you supply. If this is your workload, the honest migration target is not another OpenAI model. It is an open-weights model you run yourself — where log-probabilities are simply available because you hold the weights — or a vendor that still ships base models. That is a different procurement conversation than the one the deprecation table implies, and it is the kind of forced re-platforming that turns a single-vendor API into a contract question rather than a config question.
The price the migration column does not show
gpt-3.5-turbo-instruct bills $1.50 per million input tokens and $2.00 per million output. gpt-5.6-terra bills $2.00 input, $0.20 cached input, $12.00 output.
Input rises 1.3x. Output rises 6x.
For gpt-3.5-turbo — shutting down on 23 October, also pointed at gpt-5.6-terra — the old rate was $0.50 in and $1.00 out. That is 4x input and 12x output.
Then there is the line item no rate card shows. gpt-5.6-terra is a reasoning model, with effort levels running none, low, medium (the default), high, xhigh and max. Reasoning tokens bill as output tokens. Take a high-volume classifier at 1,000 input and 50 output tokens, one million calls a month:
| Configuration | Monthly cost | vs. instruct |
|---|---|---|
gpt-3.5-turbo-instruct | ~$1,600 | — |
gpt-5.6-terra, effort none | ~$2,600 | 1.6x |
gpt-5.6-terra, default medium (assume 300 reasoning tokens/call) | ~$6,200 | 3.9x |
The reasoning-token figure is an assumption, not a published number, and your workload will differ. The point survives the uncertainty: the gap between the 1.6x outcome and the 3.9x outcome is one parameter, and a migration script that only swaps the model string will not set it. This is the same arithmetic trap as the 272K context cliff on GPT-6 Astra — the sticker rate is not the bill.
Two genuine compensations, which may or may not reach you. terra has a 1,050,000-token context window against instruct’s 4,000, and cached input at $0.20 is a real 10x discount on repeated prefixes. Neither does anything for a short-prompt, high-volume job. And if cost is the binding constraint, note that GPT-6 Luna landed on 22 September at $0.10 in / $0.50 out — cheaper on output than the model you are being migrated to, and cheaper than the one you are leaving. OpenAI’s recommended replacement is not its cheapest option, and nothing in the deprecation table mentions that. One caveat if you are migrating an image-bearing workload onto the GPT-6 tier this week: OpenAI fixed an image-encoding defect in both Sol and Luna on 25 September, three days after they launched, and neither model exposes a dated snapshot to pin. Benchmark the vision path yourself rather than trusting a figure published before Thursday.
The contradiction in the fine-tuning rows
This is the part worth escalating internally today.
OpenAI’s deprecations page states a general rule about custom models:
“Inference on fine-tuned models will continue to be available until the base models are deprecated.”
The base models babbage-002 and davinci-002 are deprecated on 28 September 2026. By that rule, every fine-tune built on them stops on 28 September.
The same page, in a separate group announced 22 April 2026, lists ft-babbage-002 and ft-davinci-002 with a shutdown date of 23 October 2026.
Both statements are on one page. They are 25 days apart and they cannot both be true.
The asymmetry in the risk is what decides your action. If the rule governs and you planned around the 23 October row, you lose production inference on Monday with no further warning. If the row governs and you planned around the rule, you migrate three and a half weeks early — wasted effort, nothing broken. So: assume Monday. If a legacy fine-tune is load-bearing, get OpenAI support to confirm in writing, and do not let a documentation conflict serve as your change-management plan.
The surrounding context makes this less surprising and more important. OpenAI is retiring fine-tuning as a product, on a schedule that has been running all year: organisations with no prior fine-tuning history lost the ability to create jobs on 7 May 2026; organisations with no fine-tuned inference in the preceding 60 days lost it on 2 July 2026; even continuously active customers can create no new jobs after 6 January 2027. On 23 October 2026 the remaining fine-tuned models — ft-gpt-3.5-turbo, ft-gpt-4, ft-gpt-4.1-nano-2025-04-14, ft-babbage-002, ft-davinci-002 — shut down together, alongside the gpt-3.5-turbo and gpt-4 base models themselves. If your differentiation lives in a custom OpenAI model, the platform is telling you, in five separate dated steps, to move it.
Google is doing the same thing without a date
While OpenAI runs a twelve-month clock in public, Google has demonstrated the alternative: remove availability without ever creating a deprecation entry.
Since 18 September 2026, the Gemini 2.5 model pages carry this notice verbatim:
“To ensure reliable performance for everyone, we are limiting access to the 2.5 models to users who have actively used them in the past.”
And, immediately after:
“These models are not deprecated and will continue to be served until further notice through the API.”
New projects are directed to Gemini 3.5 Flash-Lite or 3.8 Flash. So gemini-2.5-pro, gemini-2.5-flash and gemini-2.5-flash-lite have no shutdown date, appear on no deprecation calendar, and are simultaneously unavailable to any new API key or project, which gets an error naming a newer model. Developers have filed reports on precisely this, including the genuinely confusing case of a circulating 16 October 2026 deprecation date for gemini-2.5-pro sitting next to a page insisting the model is not deprecated.
For a buyer this is worse than a dated shutdown, for one structural reason: it is invisible to every tool you use to manage the risk. Deprecation trackers, SBOM and dependency scanners, vendor-risk registers and procurement checklists all key off announced end-of-life dates. A capacity-driven access gate generates none. Your existing project keeps working. CI keeps passing. The failure arrives the day somebody provisions a fresh key — a new staging environment, a new customer tenant, a disaster-recovery rebuild, a contractor’s sandbox — and it arrives with no notice period at all, because there was never a notice.
What to do before Monday
- Grep for the strings. All four dying model IDs, plus any call to
/v1/completionsand any client-library method namedcompletionsrather thanchat. Check cron jobs, notebooks, internal admin scripts and services nobody has deployed in a year — legacy model strings outlive the applications that were built around them. Developers inheriting a codebase should assume at least one is in there. - Inventory your fine-tunes and date them pessimistically. For each custom model, identify the base. Treat anything on
babbage-002ordavinci-002as failing 28 September, not 23 October. - Set reasoning effort explicitly. Anything moving to
gpt-5.6-terrainheritsmediumby default. Decide the value, write it in the call, and re-run the cost model on a $12.00 output rate before the invoice does it for you. - Route the log-probability workloads elsewhere. Anything needing
echo,suffix,best_ofor raw base-model likelihoods has no OpenAI destination after Monday. Price open-weights self-hosting or a vendor that still exposes base models — the comparison between ChatGPT’s platform and alternatives like Claude or Cohere now turns on which primitives each one still offers, not just on benchmark scores. If the work runs inside a coding workflow, the same question applies to your agent and IDE tooling, which may be pinning models you never chose. - Add a new-key canary. Once a month, provision a genuinely new API key in a genuinely new project and call every model your stack depends on. Alert on any 404. This is a twenty-line job and it is the only check that catches a soft access gate. It would have caught the Gemini 2.5 change on 18 September; no deprecation calendar would have.
The thesis, stated plainly
A deprecation page is a vendor’s account of which models are ending. It is not an inventory of your availability risk, and this week demonstrates both ways it falls short. OpenAI’s page is complete, honest, generously dated — and it still describes an API-surface removal as a model swap, points a $2.00-output workload at a $12.00-output model without comment, silently drops four parameters that have no successor, and contradicts itself by 25 days on when custom models die. Google’s page for Gemini 2.5 removes availability for every new project while stating, accurately, that nothing is deprecated.
Two vendors, two failure modes, one conclusion: the artefact you need is a test that provisions a new key and calls your models, not a calendar somebody else maintains.
Frequently asked questions
What exactly shuts down on 28 September 2026?
Four model IDs: gpt-3.5-turbo-instruct, davinci-002, babbage-002 and gpt-3.5-turbo-1106. All four are listed on OpenAI's deprecations page with an announcement date of 26 September 2025 and a shutdown date of 28 September 2026, and all four carry the same recommended replacement, gpt-5.6-terra. That is twelve months and two days of notice, which is a fair clock and longer than most of the industry gives. The detail that the table does not surface is that three of those four — gpt-3.5-turbo-instruct, davinci-002 and babbage-002 — are the entire set of models that OpenAI's legacy /v1/completions endpoint serves. The fourth, gpt-3.5-turbo-1106, is a chat-era snapshot and is the only one of the group for which 'migrate to gpt-5.6-terra' describes something close to a drop-in change. For the other three, the endpoint they live on is what is being removed, and no model is arriving to replace them on it.
Is the /v1/completions endpoint itself being deleted?
OpenAI has not announced the deletion of the endpoint, and its API reference page for Create completion does not carry a deprecation banner. What happens on 28 September is narrower and stranger: the endpoint continues to be documented while having no supported models left to call. Functionally the distinction does not help anyone. A request to /v1/completions after Monday needs a model parameter, and every value the reference page lists is shut down. Whether OpenAI eventually returns a 404 on the route or a model-not-found error on the body, the workload stops either way. Plan against the models, not against the endpoint's documentation status — this is the same lesson as the Assistants API sunset, where the interesting question was never whether the route survived but where the work moved to. Anyone maintaining a client library with a completions() method should treat it as dead code after Monday rather than waiting for a formal endpoint deprecation notice that may never come.
Which capabilities have no replacement at all?
Four parameters exist only on the legacy completions endpoint, and each one anchors a real workload. echo returns the prompt alongside the completion, which is the mechanism behind log-likelihood and perplexity scoring — you echo the text back and read the per-token probabilities. suffix, supported only on gpt-3.5-turbo-instruct, does fill-in-the-middle insertion between a prefix and a suffix. best_of generates several completions server-side and returns the highest-log-probability one. And logprobs on this endpoint returns log probabilities for up to five likely tokens per position. Combine echo with logprobs on a base model and you have the standard way to score arbitrary text under an OpenAI model, which is how a large amount of academic evaluation work, calibration research, classifier confidence scoring and prompt-selection tooling was built. davinci-002 and babbage-002 were the last OpenAI base models — not instruction-tuned, not chat-wrapped — exposing that. gpt-5.6-terra is a reasoning model on the Responses and Chat Completions APIs. It does not do raw continuation, it does not accept echo or suffix or best_of, and it will not give you a clean likelihood for a string you hand it. If that is your workload, the migration is not to gpt-5.6-terra. It is to an open-weights model you host, or to a provider that still exposes base models.
How much more expensive is the recommended replacement?
More than the table implies, and the gap widens the cheaper your old model was. gpt-3.5-turbo-instruct bills $1.50 per million input tokens and $2.00 per million output. gpt-5.6-terra bills $2.00 input, $0.20 cached input and $12.00 output. Input is 1.3x; output is 6x. For gpt-3.5-turbo, which shuts down on 23 October and also points at gpt-5.6-terra, the old rate was $0.50 in and $1.00 out — so 4x input and 12x output. Then there is the part no rate card shows. gpt-5.6-terra is a reasoning model with effort levels from none through low, medium (the default), high, xhigh and max, and reasoning tokens bill as output tokens. Take a classification job at 1,000 input and 50 output tokens, one million calls a month: gpt-3.5-turbo-instruct costs about $1,600, gpt-5.6-terra with effort set to none costs about $2,600, and the same job left on the default medium effort — assume a conservative 300 reasoning tokens per call — costs about $6,200. That is roughly 3.9x, and the difference between the 1.6x and the 3.9x version is a single parameter most migration scripts will not set. The compensations are real but may not apply to you: terra has a 1,050,000-token context window against instruct's 4,000, and cached input at $0.20 is a genuine 10x discount on repeated prefixes. Neither helps a short-prompt, high-volume classifier.
When do fine-tunes built on davinci-002 and babbage-002 actually stop working?
OpenAI's documentation gives two answers and they are 25 days apart. The stated general rule is that 'inference on fine-tuned models will continue to be available until the base models are deprecated.' The base models babbage-002 and davinci-002 are deprecated on 28 September 2026. Applying the rule, the fine-tunes stop on 28 September. But the same deprecations page lists ft-babbage-002 and ft-davinci-002 in a separate group, announced 22 April 2026, with a shutdown date of 23 October 2026 and gpt-5.6-terra as the replacement. Both claims are on one page and they are not compatible. There is no ambiguity about the direction of risk: if the rule governs, a team reading the 23 October row loses production inference on Monday with no warning. If the row governs, a team reading the rule migrates three and a half weeks early, which costs effort but breaks nothing. Assume Monday, verify with OpenAI support in writing if a legacy fine-tune is load-bearing, and do not let a documentation conflict become your change-management plan. The wider context is that OpenAI is winding fine-tuning down entirely: organisations with no fine-tuning history lost the ability to create jobs on 7 May 2026, organisations with no fine-tuned inference in 60 days lost it on 2 July 2026, and even active customers can no longer create new jobs after 6 January 2027.
Is Google doing the same thing to Gemini 2.5?
The same thing, by a different mechanism, with less warning and no date. Since 18 September 2026 the Gemini 2.5 model pages carry this notice: 'To ensure reliable performance for everyone, we are limiting access to the 2.5 models to users who have actively used them in the past.' It adds that 'these models are not deprecated and will continue to be served until further notice through the API,' and points new projects at Gemini 3.5 Flash-Lite or 3.8 Flash. So gemini-2.5-pro, gemini-2.5-flash and gemini-2.5-flash-lite have no shutdown date and appear on no deprecation calendar, yet a new API key or a new Google Cloud project calling them gets an error naming a newer model instead. Developers have filed reports about exactly this, including the confusing case of a published deprecation date for gemini-2.5-pro of 16 October 2026 sitting alongside a page that says the model is not deprecated. The practical effect is worse than a dated shutdown for one specific reason: it is invisible to the tooling buyers actually use. Deprecation trackers, dependency scanners and procurement checklists all key off announced end-of-life dates. A capacity-driven access gate produces none, so your existing project keeps working, CI keeps passing, and the failure surfaces the day someone provisions a fresh key — a new environment, a new customer tenant, a disaster-recovery rebuild.
What should a buyer or engineering lead do this week?
Five things, in order. First, grep your codebase and your vendor SDKs for the four dying model IDs plus any call to /v1/completions or a client method named completions rather than chat — legacy strings survive in cron jobs, notebooks, internal scripts and abandoned microservices far longer than in the main application. Second, list every fine-tuned model you own, identify its base, and treat any legacy fine-tune as failing Monday rather than on 23 October. Third, for anything moving to gpt-5.6-terra, set reasoning effort explicitly instead of inheriting the medium default, and re-run your cost model on the new output rate before the invoice does it for you. Fourth, separate out the workloads that need echo, suffix, best_of or raw log-probabilities and route them somewhere else entirely, because no OpenAI model will serve them after Monday. Fifth, and this is the durable change: stop treating a vendor deprecation page as a complete inventory of your availability risk. Add a standing check that provisions a genuinely new API key against every model your stack depends on, monthly, and alerts when one of them 404s. That test would have caught the Gemini 2.5 gate on 18 September. Nothing on a deprecation calendar would have.
Sources
- OpenAI — Deprecations (primary: shutdown dates, replacements, fine-tuning timeline)
- OpenAI API Reference — Create completion (primary: the three models the legacy endpoint serves, and its unique parameters)
- OpenAI — GPT-5.6 Terra model page (primary: context window, endpoints, reasoning effort levels)
- OpenAI — API changelog (primary: GPT-6 Sol and Luna, 22 September 2026)
- OpenAI — GPT-4 API general availability and deprecation of older models in the Completions API
- Google — Gemini 2.5 Pro model page (primary: the access limitation notice, verbatim)
- Google — Gemini API release notes (primary: 18 September 2026 access change)
- Google AI Developers Forum — gemini-2.5-pro returns 'no longer available to new users', contradicting the published deprecation date
- PricePerToken — gpt-3.5-turbo-instruct rate card
- PricePerToken — gpt-5.6-terra rate card
- OpenAI Developer Community — Deprecation notice: upcoming model shutdowns in 2026
- OpenAI Developer Community — Will the Completion endpoint be dropped?
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.