What OpenAI turns off on October 23, 2026, how to find every call that will fail, which replacement to pick, and the code changes a model swap does not cover
On October 23, 2026, OpenAI switches off 17 models and fine-tune families in one day, including gpt-4, gpt-3.5-turbo and o1.
OpenAI's deprecations page lists the batch in a single table: gpt-4, gpt-4-turbo, gpt-3.5-turbo, gpt-4.1-nano, gpt-4o-2024-05-13, o1, o1-pro, o3-mini, o4-mini and gpt-image-1, plus every fine-tuned model built on gpt-4, gpt-3.5-turbo, gpt-4.1-nano, o4-mini, babbage-002 and davinci-002 - OpenAI deprecations. The notice went out by email on April 22, which gave developers 184 days to move. The GPT-4 deprecation is the headline, but it is the start of a run: Agent Builder, the Evals API and the v1/prompts endpoint shut down on November 30, and the original GPT-5 and o3 snapshots follow on December 11.
The hard part is not the model string you remember writing. As of October 11, the default ChatOpenAI model in both the Python and JavaScript versions of LangChain is still gpt-3.5-turbo, and so is the default model in LlamaIndex's OpenAI integration and several Instructor helpers - LangChain.js issue #11891. An app that never named a model at all will start failing on October 23 because a library named one for it. And the replacements OpenAI lists point to its GPT-5.6 family from July, not to the GPT-6 models it shipped in September, so the "official" swap is sometimes the more expensive one: OpenAI's substitute for gpt-3.5-turbo costs 8x more per output token, while a newer model in the same catalog costs a third or less of what gpt-3.5-turbo did.
This guide covers exactly what shuts down and when, why OpenAI retires models at this pace, how to find every call (including the ones hiding in dependencies, config and databases), how to pick a replacement per workload, the breaking changes that a model-string swap does not fix, the special cases (fine-tunes, images, o1-pro), the November and December deadlines, the real cost math, how Anthropic, Google, Azure and OpenRouter handle the same problem differently, a runbook for the twelve days left, and how to build so the next retirement is a config change rather than an incident.
Contents
- What Shuts Down on October 23, and Everything After It
- Why OpenAI Retires Models This Often
- Find Every Call Before It Fails
- The Code You Did Not Write: Library and Framework Defaults
- Picking the Replacement: OpenAI's Map vs Its Own Newer Models
- The Breaking Changes a Model Swap Does Not Fix
- Special Cases: Fine-Tunes, Images, o1-pro and Batch Jobs
- November 30 and December 11: Agent Builder, Evals, Prompts and GPT-5
- The Cost Math of Migrating
- How Other Providers Retire Models, and What That Teaches
- A Migration Runbook for the Next Twelve Days
- Design So the Next Retirement Is a Config Change
- Conclusion: Where to Land
Where to Land: Replacement Options Scored
Before the details, here is the short answer to the question most people arrive with: what should a GPT-4-era workload run on after October 23? The table scores ten realistic landing spots, including OpenAI's own listed substitutes (the GPT-5.6 family), the newer GPT-6 models, two Claude models, a Gemini model and a self-hosted open-weight model. Prices are list prices per million tokens on October 11, 2026, from each provider's pricing page, and the capability column uses the Artificial Analysis Intelligence Index values published in OpenRouter's model API - OpenRouter models API. That index is one lens, not a verdict on every task, and for reference the retiring gpt-4 scores 6.7 on it.
The scores assume the most common starting point: an app written against Chat Completions in the GPT-4 era, sending temperature, max_tokens and sometimes function definitions. If your starting point is different (an image pipeline, a fine-tune, an agent built in Agent Builder), the special-case sections later in this guide matter more than the table.
| # | Option | What It Does | Price (30%) | Migration effort (20%) | Capability (20%) | Shelf life (20%) | Control (10%) | Final |
|---|---|---|---|---|---|---|---|---|
| 1 | gpt-6-luna | OpenAI's newest efficient tier, also powers the Decisions API | 10 - $0.10 in / $0.50 out | 8 - supports none effort, Chat Completions works, tools need none | 6 - index 38.1 | 8 - released Sep 22, 2026, 6-month notice minimum | 4 - one vendor | 7.8 |
| 2 | Claude Haiku 5.5 | Anthropic's small model, same price as gpt-6-luna under 100K-token prompts | 10 - $0.10 / $0.50 (5x above 100K, which no GPT-3.5 prompt reaches) | 3 - new SDK and request shape, prompts need re-testing | 7 - index 43.4 | 9 - "not sooner than Oct 7, 2027" | 7 - second vendor, served on several platforms Anthropic operates | 7.5 |
| 3 | Claude Sonnet 5.5 | Anthropic's mid tier, highest index in this table | 7 - $2 / $10 | 3 - new SDK and request shape | 10 - index 56.0 | 9 - "not sooner than Sep 28, 2027" | 7 - second vendor, several Anthropic-operated platforms | 7.2 |
| 4 | gpt-5.6-luna | OpenAI's listed substitute for gpt-4.1-nano | 9 - $0.20 / $1.20 | 8 - same shape as GPT-5.6 Sol | 6 - index 37.3 | 6 - GPT-5.x line already being retired | 4 - one vendor | 7.1 |
| 5 | gpt-6.1-sol | Near-Astra model launched at DevDay, Sep 29 | 7 - $2 / $10 | 5 - no none effort, tools require Responses API | 9 - index 51.8 | 8 - newest release, no deprecation | 4 - one vendor | 6.9 |
| 6 | gpt-5.6-terra | OpenAI's listed substitute for gpt-3.5-turbo and o4-mini | 6 - $2 / $12 | 8 - Chat Completions with none effort | 7 - index 42.1 | 6 - GPT-5.x line already being retired | 4 - one vendor | 6.4 |
| 7 | gpt-5.6-sol | OpenAI's listed substitute for gpt-4, gpt-4-turbo and o1 | 5 - $4 / $20 (promo price through at least Nov 21) | 8 - Chat Completions with none effort | 8 - index 47.0 | 6 - GPT-5.x line already being retired | 4 - one vendor | 6.3 |
| 8 | gpt-oss-120b | OpenAI's open-weight model, Apache 2.0, runs on one H100 | 8 - $0.037 / $0.17 via hosts, or your own GPU | 2 - you run or rent the serving stack | 2 - index 11.6 (still above gpt-4's 6.7) | 10 - weights never retire | 10 - fully yours | 6.2 |
| 9 | Gemini 3.8 Flash | Google's newest Flash model | 7 - $0.75 / $3.75, doubles Jan 1, 2027 | 3 - new SDK and request shape | 7 - index 40.9 | 7 - no shutdown date, dates are "earliest possible" | 5 - second vendor, one company's cloud | 6.0 |
| 10 | gpt-6-astra | OpenAI's flagship, per its models page | 3 - $10 / $50 | 4 - no none, no temperature, tools need Responses, asks more questions | 9 - index 52.7 | 8 - released Sep 3, 2026 | 4 - one vendor | 5.5 |
How the criteria work. Price (30%) scores list price per million tokens, mostly output, because output is where reasoning tokens land. Migration effort (20%) scores how much GPT-4-era Chat Completions code has to change. Capability (20%) maps the index into bands (55 and up is a 10, each five points lower drops one). Shelf life (20%) scores how long you can expect to stay before the next forced move, using each provider's published dates. Control (10%) scores how independent the option leaves you: a second vendor beside OpenAI scores higher, and open weights you can run anywhere score highest.
Two things stand out. First, OpenAI's own listed substitutes land in the middle of the table, behind a newer OpenAI model that is cheaper per token. Second, every OpenAI-hosted option scores 4 on control, because a model that only one company serves is a model whose retirement date that company sets. Neither point means you should leave OpenAI: it means the replacement deserves a decision, not a copy-paste from the deprecations table.
1. What Shuts Down on October 23, and Everything After It
The October 23 batch is the largest single shutdown OpenAI has scheduled in 2026, and it is not a cleanup of obscure snapshots. It removes the original GPT-4, the model most tutorials, courses and early AI products were written against, along with gpt-3.5-turbo, which was the default "cheap model" in a generation of code samples. It also removes the first two generations of OpenAI's reasoning models (o1, o1-pro, o3-mini, o4-mini) and the first GPT Image model. OpenAI's stated reason, in the email it sent on April 22, was "to improve reliability and simplify model selection" - OpenAI developer forum.
What makes this batch different from earlier ones is who is still on it. Earlier shutdowns mostly hit preview models and dated snapshots that people pinned on purpose. This one hits bare aliases: gpt-4, gpt-3.5-turbo, o1, o4-mini. Those are the names people type when they do not think about versions, which is exactly the code that nobody revisits. The forum post announcing the batch has more than 7,000 views, and the first replies describe it as an unprecedented list and object most strongly to losing fine-tuned models that teams spent months tuning.
The October 23 list, with OpenAI's substitutes
OpenAI's table maps each retiring snapshot to a substitute. Every alias in the left column dies with its snapshot, so gpt-4 stops working along with gpt-4-0613, and o1 along with o1-2024-12-17. The prices are standard per-million-token rates from OpenAI's pricing page on October 11 - OpenAI pricing.
| Retiring model (and aliases) | Price in / out | OpenAI's substitute | Substitute price in / out |
|---|---|---|---|
gpt-4-0613 (gpt-4) | $30 / $60 | gpt-5.6-sol | $4 / $20 |
gpt-4-turbo (gpt-4-turbo-2024-04-09) | $10 / $30 | gpt-5.6-sol | $4 / $20 |
gpt-4o-2024-05-13 | $5 / $15 | gpt-5.6-sol | $4 / $20 |
gpt-3.5-turbo-0125 (gpt-3.5-turbo) | $0.50 / $1.50 | gpt-5.6-terra | $2 / $12 |
gpt-4.1-nano | $0.10 / $0.40 | gpt-5.6-luna | $0.20 / $1.20 |
o1 | $15 / $60 | gpt-5.6-sol | $4 / $20 |
o1-pro | $150 / $600 | gpt-5.6-sol with reasoning.mode: pro | $4 / $20 plus pro work |
o3-mini | $1.10 / $4.40 | gpt-5.6-sol | $4 / $20 |
o4-mini | $1.10 / $4.40 | gpt-5.6-terra | $2 / $12 |
gpt-image-1 | $10 / $40 (image tokens) | gpt-image-2.5-sunburst or -flare | $8 / $30 (image tokens) |
Two details in the table deserve attention before you trust it. The gpt-4-1106-preview row appears in the October 23 list even though the same page records that snapshot as already shut down on March 26, 2026, and a developer pointed out in April that the email named gpt-4-turbo without its dated snapshot (today's page lists gpt-4-turbo-2024-04-09 as an alias in the row). Small inconsistencies like these are a reason to verify by calling the API after a migration rather than by reading the table. The fine-tune rows map ft-gpt-4 to gpt-5.6-sol, ft-gpt-3.5-turbo to gpt-5.6-terra and ft-gpt-4.1-nano to gpt-5.6-luna, but none of those substitutes can be fine-tuned, which section 7 covers.
The rest of the calendar
October 23 is the first of five dates that matter to anyone building on OpenAI this quarter. Each one hits a different layer of the stack: models, then managed products, then the newer GPT-5 snapshots that many teams moved to only last year. Treating them as one migration project, rather than five fire drills, is the cheapest way through.
| Date | What shuts down | Recommended path |
|---|---|---|
| Oct 23, 2026 | The batch above, plus its fine-tunes | GPT-5.6 or GPT-6 models, GPT Image 2.5 |
| Oct 31, 2026 | Existing evals become read-only | Recreate evals in Promptfoo or your own scripts |
| Nov 30, 2026 | Agent Builder, the Evals dashboard and API, v1/prompts | Agents SDK or Workspace Agents, Promptfoo, prompts in code |
| Dec 1, 2026 | gpt-image-1-mini, gpt-image-1.5, chatgpt-image-latest | GPT Image 2.5 Sunburst or Flare |
| Dec 11, 2026 | gpt-5-2025-08-07, gpt-5-mini, gpt-5-nano, gpt-5-pro, o3, o3-pro snapshots | gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna |
Further out, OpenAI has already scheduled gpt-5.1, gpt-5.3-codex and gpt-5.4-nano for April 1, 2027, with gpt-6-sol and gpt-6-luna as the replacements, and whisper-1 plus the GPT-4o transcription models for February 26, 2027. The Assistants API already shut down on August 26, 2026, so any app still calling it has been broken for six weeks. The pace is visible when you count the rows on the deprecations page by the quarter in which each shutdown lands.
The jump in the second half of 2026 is the point. A team that migrated once a year in 2024 now faces a forced change every few weeks somewhere in its OpenAI footprint, and the fourth quarter of 2026 alone accounts for 29 shutdowns. Why this matters: the cost of a migration is mostly fixed (finding the calls, testing, deploying), so a higher frequency multiplies that fixed cost. How to apply it: budget for retirements as a recurring maintenance line, the way you budget for dependency upgrades, and build the seam described in section 12 so each one gets cheaper.
2. Why OpenAI Retires Models This Often
It is tempting to read a deprecation wave as a product decision that could have gone the other way, or as a push toward pricier models. The structural reason is simpler and it applies to every lab. A served model is not a file on a disk: it is a fleet of GPUs with that model's weights loaded, kept warm, load-balanced, monitored and patched. Anthropic said it plainly when it explained its own retirement policy: "the cost and complexity to keep models available publicly for inference scales roughly linearly with the number of models we serve" - Anthropic, Commitments on model deprecation.
Run that forward. Each new model a lab ships adds a fleet. The GPUs that serve gpt-4 to a shrinking set of customers at $30 per million input tokens are GPUs that cannot serve a model that customers are lining up for. When usage of the old model falls below the cost of its share of capacity, retiring it is close to forced, and the faster new models arrive, the faster that crossover comes. OpenAI shipped GPT-5.6 in July, GPT-6 Astra on September 3, GPT-6 Sol and Luna on September 22 and GPT-6.1 Sol on September 29 - OpenAI API changelog. Four model families in under three months is four new fleets competing for the same chips.
The notice policy, and what it does and does not promise
OpenAI does publish rules for how much warning it gives, and it is worth knowing them precisely, because they set your planning horizon. Generally available models get at least six months of notice, "specialized variants" such as chat, Codex and deep research models get at least three months, and anything with preview in its name may be retired "with much shorter notice, such as 2 weeks" - OpenAI deprecations. Safety or compliance concerns can shorten any of these.
The same page separates two words that people use loosely. Legacy means a model no longer receives updates and will probably be deprecated later. Deprecated means a shutdown date exists. Once that date passes, the model "will no longer be accessible." The page also mentions one escape hatch that few people notice: "In some cases, developers may be able to provision dedicated capacity for continued access after a model's shutdown date," through OpenAI's sales team. That is an option for an enterprise with a regulated workload that cannot change on schedule, not for a startup, and it confirms the point above: continued access to an old model is a capacity purchase.
The pattern is not limited to models, either. The same week this batch approaches, Vercel announced that starting October 23 it deletes deployments older than 30 days for Pro teams unless they opt out, and starts billing deployment storage at $0.10 per GB-month - Vercel changelog. Cloudflare announced it is absorbing the Deno team, and Deno Deploy will shut down in six months - The New Stack. Platforms reclaim the resources that old things hold, on their schedule. If your app runs on Vercel and calls gpt-4, October 23 is a two-deadline day.
Why this matters: once you see retirements as a capacity decision, you stop hoping for an extension and start planning for a cadence. How to apply it: assume every hosted model you use today has a shutdown date within 12 to 24 months, check the deprecation pages of every provider you call once a month, and favor models released recently, because they sit at the far end of that window.
3. Find Every Call Before It Fails
The failure you are trying to prevent is specific and easy to recognize once it happens. After a shutdown, a request that names the retired model returns an HTTP 404 with the error code model_not_found and a message of the form "The model ... does not exist or you do not have access to it" - OpenAI developer forum. OpenAI's error guide files this under NotFoundError, with the advice to check the resource identifier - OpenAI error codes. If your code treats every 4xx as a user error and shows a generic message, the outage looks like a bug in your product rather than a vendor change.
There is a second trap. After the July 23 batch, a developer on the OpenAI forum tested the retired models and found that most were "still listed by the models endpoint - but have been disabled": Chat Completions answered "The model ... has been deprecated" while the Responses API answered "Model not found" - OpenAI developer forum. That is a community test rather than documented behavior, but it is enough to rule out one tempting shortcut: checking GET /v1/models at startup does not prove a model still serves requests. Only a real request does.
Where model IDs hide
Model identifiers end up in more places than application code. The diagram below shows where to look, grouped by who wrote the string. The last two groups are the ones teams miss, because no grep of your own src/ folder finds them.
The second group is the one that bites products with customers. If your app ever let a user choose a model, or stored the model alongside a saved agent, a document template or a workspace setting, the retiring IDs are sitting in your database. Changing the default in code does nothing for those rows. The same is true of Batch API input files written before the shutdown and of scheduled jobs whose payloads were serialized weeks ago.
Three searches that cover most of it
The fastest audit combines a code search, a usage query and a log query. Each catches something the others miss, so run all three. Start with the code search, which takes seconds and needs nothing but a checkout of every repository that talks to OpenAI.
# Every retiring ID, as a quoted string, across code, config and templates.
# Run it in each repository, including infrastructure and data repos.
rg -n --hidden -g '!node_modules' -g '!.git' \
-e " ['\"](gpt-4|gpt-4-0613|gpt-4-turbo|gpt-4-turbo-2024-04-09|gpt-4-1106-preview) ['\"]" \
-e " ['\"](gpt-3\.5-turbo|gpt-3\.5-turbo-0125|gpt-4\.1-nano|gpt-4o-2024-05-13) ['\"]" \
-e " ['\"](o1|o1-pro|o3-mini|o4-mini|gpt-image-1) ['\"]" \
-e "ft:(gpt-4|gpt-3\.5-turbo|gpt-4\.1-nano|o4-mini|babbage-002|davinci-002)"
The quoted-string pattern avoids false positives such as gpt-4o or gpt-4.1 (which are not on this list), and the last line catches fine-tuned model IDs, which start with ft: followed by the base model. Next, ask OpenAI what your organization actually called. The Usage API returns completions usage grouped by model and project, which shows traffic that no code search will, such as a forgotten cron job or a teammate's script. It needs an admin key from your organization settings - OpenAI Cookbook, Usage API.
# Last 30 days of completions usage, grouped by model and project.
curl -s -G "https://api.openai.com/v1/organization/usage/completions" \
-H "Authorization: Bearer $OPENAI_ADMIN_KEY" \
--data-urlencode "start_time=$(( $(date +%s) - 30*24*3600 ))" \
--data-urlencode "bucket_width=1d" \
--data-urlencode "limit=31" \
--data-urlencode "group_by=model" \
--data-urlencode "group_by=project_id" \
| jq ' [.data [].results []] | group_by(.model)
| map({model: . [0].model, requests: (map(.num_model_requests) | add)})'
Any model in that output that appears in the October 23 table is a live dependency. Drop the final jq step to see the raw results, where each row also carries the project_id, which tells you which team owns the traffic. Finally, search your own request logs for the model field. If you log outgoing requests, the model name is in them; if you do not, this deprecation is a good reason to start, because the same log answers "what changed" during the migration.
Together, the three searches answer the questions that decide your plan. The code search says where the strings live, the usage query says which of them carry real traffic, and the logs show which users and features sit behind that traffic. Why this matters: teams that skip the usage query routinely find a second integration after the shutdown, usually owned by someone who left. How to apply it: write the results into one table (model, location, owner, monthly tokens, replacement), because that table becomes the migration plan, the test plan and the rollback plan.
4. The Code You Did Not Write: Library and Framework Defaults
The most widespread source of October 23 failures will be code that never mentions a model, because a framework fills in the default. In the JavaScript version of LangChain, new ChatOpenAI() with no model sends gpt-3.5-turbo: the class still declares model = "gpt-3.5-turbo" on its main branch - LangChain.js source. The Python package declares the same default with Field(default="gpt-3.5-turbo", alias="model") - langchain-openai source. LlamaIndex's OpenAI integration sets DEFAULT_OPENAI_MODEL = "gpt-3.5-turbo" - LlamaIndex source.
These are not hypothetical reports. Developers filed issues against each project between October 2 and October 8. The LangChain.js report notes that two other defaults in the same package already point at dead models: the completions class defaults to gpt-3.5-turbo-instruct, shut down September 28, and the DALL-E wrapper defaults to dall-e-3, shut down May 12 - LangChain.js issue #11891. A pull request in the Python repository proposes the cleanest fix, which is to require an explicit model and stop guessing - langchain PR #41041. It is still open as of this writing, and the Python issue it answers is open too - langchain issue #40995.
Who else is affected
The pattern repeats across the ecosystem because it was the sensible default in 2023: gpt-3.5-turbo was cheap, fast and available to everyone, so libraries reached for it when the caller did not say. The Instructor library uses it as the default for its LLM validator and its distillation helpers, and its maintainers were told that the validator also sends temperature=0, which may not suit a reasoning model - Instructor issue #2735. The lagent agent framework hardcodes both gpt-3.5-turbo and gpt-4o-2024-05-13, and the report on it was found by a developer running a scanner over public repositories for upcoming platform deadlines - lagent issue #378.
What makes defaults dangerous is that they fail at a distance. The error surfaces in your logs as a 404 from OpenAI, with a stack trace that points into the library rather than into your code, and nothing in your repository mentions the model it complains about. Teams that grep their own source for gpt-3.5-turbo, find nothing and conclude they are safe are exactly the teams this catches. The four public cases with dated evidence are below.
- LangChain (JS and Python):
ChatOpenAI()defaults togpt-3.5-turbo - LlamaIndex: the OpenAI LLM class defaults to
gpt-3.5-turbo - Instructor: validator and distillation helpers default to
gpt-3.5-turbo - Smaller frameworks:
lagentandx-crawlhardcode retiring IDs - x-crawl issue #183
The list is short on purpose: these are the cases with public, dated evidence. The real exposure is wider, because every wrapper, starter template and internal SDK built on top of those libraries inherits the default, and because AI coding assistants learned from code written when gpt-3.5-turbo and gpt-4 were the standard choices. A coding model whose training data is mostly 2023 and 2024 code can easily write model="gpt-4" into a new file today, because that is what most of the examples it learned from say. If any part of your app was generated by an AI builder or a coding assistant and never reviewed line by line, treat it as a likely carrier, which is one of the reasons we argued for a deliberate hand-off in when to graduate from a vibe-coding tool.
The fix: never rely on a default model
The repair is the same in every framework: name the model explicitly, and read it from one place. Do not wait for the library to change its default, because when it does, it will pick a replacement with its own price and behavior, and you will inherit that silently.
# Python: LangChain and LlamaIndex, explicit model from one setting
import os
from langchain_openai import ChatOpenAI
from llama_index.llms.openai import OpenAI as LlamaOpenAI
CHAT_MODEL = os.environ ["APP_CHAT_MODEL"] # for example "gpt-6-luna"
chat = ChatOpenAI(model=CHAT_MODEL) # never ChatOpenAI() bare
llm = LlamaOpenAI(model=CHAT_MODEL) # never OpenAI() bare
// TypeScript: LangChain.js, same rule
import { ChatOpenAI } from "@langchain/openai";
const CHAT_MODEL = process.env.APP_CHAT_MODEL!; // fail fast if unset
export const chat = new ChatOpenAI({ model: CHAT_MODEL });
Reading the model from an environment variable is not about flexibility for its own sake. It makes the model a deployment setting, so the next retirement is a config change and a redeploy rather than a code change, a review and a release. Make the app refuse to start when the variable is missing, so a forgotten setting fails at deploy time instead of falling back to a library default that may be dead.
The same rule applies to the wrappers you wrote yourself. Many codebases have a small llm.py or ai.ts helper with a signature like complete(prompt, model="gpt-4"), written once and called from dozens of places that never pass the model. Those helper defaults are the in-house version of the library problem, and they are worth hunting for specifically: search for function signatures whose default argument is a model name, and change the default to the shared setting.
Finally, check the places where a model name arrives from outside your code at runtime. Admin panels, per-customer settings and feature configuration stored in a database can all carry a model string that a library or helper then passes straight through. Section 3's audit catches those values in the data; this step makes sure the code that reads them validates against a list of models you actually support, and rejects or maps anything else instead of sending it to OpenAI.
Why this matters: a default model is a dependency you did not choose, with a shutdown date you were not told about. How to apply it: add one more search to the audit in section 3 that finds bare constructors (ChatOpenAI(), OpenAI() from LlamaIndex, Instructor's llm_validator without model=), fix them to read the shared setting, and pin your framework versions so a future default change arrives through a reviewed upgrade.
5. Picking the Replacement: OpenAI's Map vs Its Own Newer Models
OpenAI's substitute column for the October 23 batch points to the GPT-5.6 family, released on July 9: Sol for the old flagships and reasoning models, Terra for the mid tier, Luna for the smallest. The April email that announced the batch named no specific replacements, only "recommended newer alternatives", so the column was filled in on or after July 9, and it has not moved since the GPT-6 releases in September, even though OpenAI's own October 1 deprecation entry for gpt-5.1 and gpt-5.4-nano already recommends gpt-6-sol and gpt-6-luna. OpenAI's model pages describe Sol as roughly "the unsuffixed model tier used in earlier GPT-5 families" and Terra as roughly "the mini model tier" - GPT-5.6 Terra model page. The map was sensible in July. Since then OpenAI has shipped four more models (GPT-6 Astra, GPT-6 Sol, GPT-6 Luna and GPT-6.1 Sol), and its models page now tells new users to start with GPT-6 Astra, choose GPT-6.1 Sol "to balance intelligence and cost", or GPT-6 Luna "for cost-sensitive, high-volume workloads" - OpenAI models.
So the deprecations page and the models page currently give different advice, and neither is wrong. The deprecations page names a substitute that is a close functional match for the old request shape. The models page names what OpenAI wants new work built on. The gap between them is where the money is, because per-token prices did not move in one direction: for the old flagships, the listed substitute is far cheaper, and for the old cheap models, it is far more expensive.
Read the chart from the left. A gpt-4 or o1 workload moving to gpt-5.6-sol pays a third of what it paid per output token, and its input price falls from $30 or $15 to $4. A gpt-3.5-turbo workload moving to gpt-5.6-terra pays 8x per output token and 4x per input token, before any reasoning tokens. The reason is that the old cheap models were small, non-reasoning models, and OpenAI's tiers no longer include anything quite that small at the Terra level. For those workloads, the better OpenAI landing spot is gpt-6-luna, at $0.10 input and $0.50 output per million tokens, which is cheaper than gpt-3.5-turbo was on both sides - GPT-6 Luna model page.
Capability is not the constraint any more
The usual worry about a forced migration is that the replacement will be worse. For this batch, the numbers say the opposite by a wide margin. On the Artificial Analysis Intelligence Index values in OpenRouter's catalog, the retiring models score between 5.5 and 16.7, and every replacement scores above 37 - OpenRouter models API.
An index is an average over benchmarks, and an average hides the tasks where a specific old model was unusually good at your specific prompt. But a gap of this size changes the migration question. You are not looking for a model as capable as gpt-4: almost anything current clears that bar. You are looking for the cheapest model that keeps your outputs stable, and the answer is often a tier or two below the substitute OpenAI lists. We walked through how to choose between OpenAI's tiers for different request types in GPT-5.6 Sol vs Terra vs Luna, and the same reasoning carries to the GPT-6 family.
OpenAI introduced GPT-6.1 Sol at DevDay on September 29, and the keynote is the most direct account of where OpenAI is steering new API work. It is long, so skip to the model segment if you only want the API changes.
The keynote frames GPT-6.1 Sol as near-Astra performance at a lower price, and OpenAI's changelog prices it at $2 input and $10 output per million tokens for prompts up to 272K tokens, which is half the per-token price of gpt-5.6-sol - OpenAI API changelog. There is one catch that the keynote does not dwell on and that matters a great deal for GPT-4-era code: GPT-6.1 Sol does not support the none reasoning effort, and tool calling requires the Responses API. Section 6 explains why that turns a one-line swap into a small rewrite.
A decision path per workload
The right replacement depends on what the old model was doing, not on which model it was. A gpt-4 call that classifies support tickets and a gpt-4 call that drafts legal summaries should not land in the same place. The diagram below is a starting point for that decision, built from the price and capability data above.
The left branch matters more than it looks. A large share of gpt-3.5-turbo and gpt-4 traffic in production is classification, routing and extraction: the model reads text and returns a category. OpenAI launched the Decisions API in beta on October 6 for exactly that shape: it returns a probability, a choice from a fixed set, or a score, using gpt-6-luna, and it bills only input tokens at $0.10 per million, with no output charges - OpenAI Decisions guide. OpenAI says it answers about 10x faster than the Responses API. It is beta and currently supports one model, so treat it as a candidate to test rather than a default, but for a pipeline that asks "which department is this ticket for?" a million times a month it is the cheapest path on OpenAI's platform.
Why this matters: picking the substitute from the deprecations table is the fastest decision and often the most expensive one. How to apply it: for each row of your audit table, write down what the call does (classify, extract, chat, reason, generate images), pick the landing model from the path above, and keep OpenAI's listed substitute as the fallback for calls you cannot change before October 23.
6. The Breaking Changes a Model Swap Does Not Fix
If you change model="gpt-4" to model="gpt-5.6-sol" and deploy, some requests will work, some will fail with HTTP 400, and some will succeed while costing more or returning empty text. The reason is that every model on the replacement side is a reasoning model, and the retiring gpt-4, gpt-4-turbo and gpt-3.5-turbo were not. Reasoning models think before they answer, they bill that thinking as output tokens, and they reject or ignore several parameters that GPT-4-era code sends by habit. This section goes through each change in the order you will hit it, with the code to fix it.
The changes below come from OpenAI's model pages, its reasoning guide and its GPT-6 migration notes. Where OpenAI documents a rule for one family but not another, the text says so, and the safe default is to test rather than assume. If you are also moving Claude workloads this quarter, the same class of problem (parameters that newer reasoning models reject) is covered in our Claude Opus 5.5 migration guide.
Change 1: Reasoning is on by default, and it is billed as output
The GPT-5.6 models default to medium reasoning effort when you do not set one, and so do GPT-6.1 Sol, GPT-6 Sol and GPT-6 Luna - OpenAI reasoning guide. Reasoning tokens are invisible in the response but "still occupy space in the model's context window and are billed as output tokens." A gpt-3.5-turbo call that returned a five-token label now spends some number of hidden tokens deciding on that label, at the output rate, on every request. For a chatbot that is a modest cost increase; for a classifier running millions of times, it can be the largest line in the migration, which section 9 quantifies.
OpenAI's own diagram shows how reasoning tokens sit inside each turn and why a tight output limit can truncate the visible answer.
The third column is the failure to watch for. When reasoning plus visible output reaches the limit you set, the API returns a response with status incomplete and reason max_output_tokens, "before any visible output tokens are produced" in the worst case, which means you paid for input and reasoning and received nothing. OpenAI recommends reserving at least 25,000 tokens for reasoning and output while you experiment. The fix has two parts: set the effort explicitly (use none where the model supports it and the task does not need thinking), and raise or remove output limits that were sized for a model that did not think.
Change 2: max_tokens is deprecated, and its replacement covers thinking
GPT-4-era Chat Completions code usually caps output with max_tokens. OpenAI's API reference marks that parameter as "deprecated in favor of max_completion_tokens" and "not compatible with o-series models" - OpenAI Chat Completions reference. In the Responses API the equivalent is max_output_tokens. In both cases the new limit counts reasoning tokens plus visible tokens, so a limit of 300 that comfortably fit a gpt-4 answer can now be consumed by thinking alone.
The fix is not to delete the limit, because the limit is also your protection against a runaway response. Rename the parameter, then size it from data: run a few hundred representative requests with a generous limit, read usage.output_tokens_details.reasoning_tokens and the visible output length from each response, and set the cap at a comfortable margin above the highest total you saw. Pair that with an explicit effort setting, because the effort level is what mostly decides how many reasoning tokens a request spends. Our guide to setting the effort dial shows how much the same task can vary in token use between effort levels.
Change 3: Sampling parameters disappear when reasoning is on
GPT-6 Astra "does not support custom temperature or top_p values or log probabilities," and OpenAI's GPT-6 migration notes say that whenever reasoning effort is not none, you should remove temperature, top_p and top_logprobs, plus logprobs in Chat Completions - Using GPT-6. Promptfoo's provider docs state the rule for GPT-6 Sol and Luna in one line: "Sampling and log-probability options are supported only with reasoning effort none" - Promptfoo OpenAI provider docs. The GPT-5.6 model pages do not spell it out, which is why the Instructor maintainers were warned that their temperature=0 validator call needs testing against gpt-5.6-terra. Code that relied on temperature=0 for repeatable output needs a different strategy anyway: structured outputs with a strict schema give you repeatable shape, and an evaluation set gives you evidence of repeatable content.
It helps to be clear about what temperature=0 was doing in the first place. On a non-reasoning model it made the single most likely next token win at every step, which gave mostly stable answers to identical prompts. A reasoning model reaches its answer through a hidden chain of steps, and OpenAI does not let you set sampling parameters while that reasoning is on, so the old knob no longer maps to the thing it used to control. The durable replacement is to define what "stable" means for your product (the same label, the same JSON fields, the same refusal behavior) and test for it directly, rather than asking the model to be deterministic.
Change 4: Tool calling in Chat Completions now depends on effort
This is the change most likely to break a working agent. OpenAI's Responses migration guide states that "starting with GPT-5.4, Chat Completions does not support tool calling with reasoning_effort values other than none" - OpenAI Responses migration guide. GPT-6 Luna and GPT-6 Sol follow the same rule, and GPT-6.1 Sol and GPT-6 Astra go further: they do not support none at all, so tool calling with them requires the Responses API.
The rule makes sense once you look at what a tool call does to a thinking model. In Chat Completions, a tool call ends the assistant's turn and the next request carries only visible messages, so whatever the model reasoned before choosing the tool is gone when it reads the result. In the Responses API, OpenAI recommends passing the reasoning items back with the function output because this "allows the model to continue its reasoning process," either through previous_response_id or by resending the output items - OpenAI reasoning guide. Read that way, allowing tools only at none in Chat Completions simply means there is no reasoning to lose. In practice there are two paths for a GPT-4 app that uses function calling.
- Path A, keep Chat Completions: move to a model that supports
none(GPT-5.6 or GPT-6 Luna/Sol), setreasoning_effort="none", keep your tool loop - Path B, move to Responses: switch the endpoint, flatten tool definitions, return results as
function_call_outputitems, then use any model and any effort
Path A is the October 23 fix, because it changes the fewest lines. Path B is the durable fix: OpenAI says Chat Completions "remains supported" but recommends Responses "for all new projects," and its own internal evals show a 3% SWE-bench gain and 40% to 80% better cache utilization with reasoning models on Responses. OpenAI said in March 2025 that it intended to support Chat Completions "indefinitely" - Simon Willison's notes on the launch. Both statements can be true while the newest models quietly require Responses for the features that matter most, which is the direction the GPT-6.1 Sol restrictions point.
Here is the same weather-tool call written three ways: the GPT-4 original, the minimal Path A fix, and the Path B rewrite.
# BEFORE: GPT-4-era Chat Completions call. Fails on Oct 23 (model gone).
resp = client.chat.completions.create(
model="gpt-4",
temperature=0,
max_tokens=300,
messages= [{"role": "system", "content": SYSTEM},
{"role": "user", "content": question}],
tools= [{"type": "function", "function": {
"name": "get_weather", "description": "Weather for a city",
"parameters": {"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"], "additionalProperties": False}}}],
)
# PATH A: same endpoint, a model that supports effort "none".
resp = client.chat.completions.create(
model="gpt-5.6-sol", # OpenAI's listed substitute for gpt-4
reasoning_effort="none", # required for tool calls in Chat Completions
max_completion_tokens=2000, # replaces max_tokens; size it with headroom
messages= [{"role": "system", "content": SYSTEM},
{"role": "user", "content": question}],
tools=TOOLS_CHAT_FORMAT, # unchanged nested {"type", "function": {...}}
)
# temperature removed: test it before relying on it with reasoning models
# PATH B: Responses API, any model, any effort.
TOOLS = [{"type": "function", "name": "get_weather",
"description": "Weather for a city",
"parameters": {"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"], "additionalProperties": False}}]
resp = client.responses.create(
model="gpt-6.1-sol",
reasoning={"effort": "low"}, # 6.1 Sol has no "none"; "low" is the floor
instructions=SYSTEM,
input=question,
tools=TOOLS,
max_output_tokens=4000,
)
for item in resp.output:
if item.type == "function_call":
result = get_weather(**json.loads(item.arguments))
resp = client.responses.create(
model="gpt-6.1-sol", reasoning={"effort": "low"},
instructions=SYSTEM, # resend: not carried by previous_response_id
previous_response_id=resp.id,
input= [{"type": "function_call_output",
"call_id": item.call_id, "output": json.dumps(result)}],
tools=TOOLS,
)
print(resp.output_text)
Three details in Path B trip people up. Tool definitions are flat in Responses (no nested function object). Tool results go back as function_call_output items matched by call_id. And previous_response_id does not carry the previous request's top-level instructions, so you resend them on every call, as OpenAI's guide notes. It also notes that all prior input tokens in a chain are still billed as input, so chaining is a convenience, not a discount. Structured outputs move too: response_format in Chat Completions becomes text.format in Responses.
Change 5: Caching and newer GPT-6 parameters
If you move straight to the GPT-6 family from GPT-5.5 or earlier, OpenAI's migration notes list one more rename: replace prompt_cache_retention with prompt_cache_options.ttl set to "30m", and review how cache writes are billed, since GPT-6 models charge cache writes at 1.25x the uncached input rate - OpenAI prompt caching guide. Cached reads, in exchange, cost a tenth or less of the input price (5% on GPT-6.1 Sol). For an app with a long, stable system prompt that is a large saving, and our prompt caching guide shows how to structure prompts so the cache actually hits.
Change 6: Behavior that no error message reports
The last group of changes never produces an error, which makes them the easiest to ship by accident. GPT-6 Astra is described by OpenAI as "more likely to ask for clarification where earlier models would make assumptions," and its guide includes prompts to make it "bias towards action" when your product needs it to finish a task rather than ask a question. It is also "more sensitive to instructions contained in skills and other files, such as AGENTS.md," and it "tends toward detailed, formatted responses," so a chatbot that used to answer in two plain sentences may start answering with headers and bullet lists unless the prompt says otherwise.
The context window changes the other way. gpt-4 had an 8,192-token window and a December 2023 knowledge cutoff; gpt-5.6-sol has 1,050,000 tokens and a February 2026 cutoff - GPT-5.6 Sol model page. Code that truncated history to fit 8K will keep working, but code that "sends everything" because the old model forced it to trim will now send far more, and prompts above 272K tokens are billed at 2x input and 1.5x output for the whole request. Retrieval code that was tuned to pick the top three chunks because nothing else fit may now do better with more context, which is worth testing, not assuming.
Why this matters: a migration that passes a smoke test can still change cost, latency and tone for every user. How to apply it: for each call, set reasoning effort explicitly, replace max_tokens, remove sampling parameters, choose Path A or Path B for tools, and run the same evaluation set against the old and new model before October 23, while the old one still answers.
7. Special Cases: Fine-Tunes, Images, o1-pro and Batch Jobs
Most of this guide assumes a plain text call. Four kinds of workload need their own plan, because the general advice either does not apply or actively misleads. Fine-tuned models lose their path forward, image pipelines change pricing structure, o1-pro users face a different cost model, and batch jobs can fail after the deadline even when the code was fixed before it. Each subsection below covers one of them.
Fine-tuned models have no direct successor
Every fine-tune built on a retiring base model shuts down with it: OpenAI's fine-tuning guide says fine-tuned models "remain available for inference until their base models are deprecated" - OpenAI supervised fine-tuning guide. The substitutes OpenAI lists for those fine-tunes are base models, and none of them can be fine-tuned: the model pages for GPT-5.6 Sol, Terra and Luna and for GPT-6.1 Sol all list the fine-tuning endpoint as "Not supported." The same guide opens with the larger fact: "OpenAI is winding down the fine-tuning platform." New organizations lost access on May 7, organizations without recent fine-tuned inference lost the ability to create jobs on July 2, and existing customers lose new job creation on January 6, 2027.
The developer reaction on OpenAI's forum is what you would expect. One user asked whether a model fine-tuned on gpt-4.1-mini "will never be used for inference" after its base is retired, and noted that supervised fine-tuning was only ever available on the GPT-4.1 variants and reinforcement fine-tuning only on o4-mini - OpenAI developer forum. Another called the wind-down a cost saving that "kind of hurts the main appeal of using OpenAI services" - OpenAI developer forum. The loss is real, but it is smaller than it looks for most teams, because many fine-tunes were compensating for a weak base model: a fine-tune of gpt-3.5-turbo built to follow a format or a tone taught a 2023 model something a 2026 model does from instructions.
The practical options, in the order most teams should try them:
- Prompt the replacement with your instructions and 10-30 examples from the training set, cached
- Use structured outputs for any fine-tune whose real job was "always return this JSON shape"
- Fine-tune elsewhere, on a provider or open-weight model you can train and keep
- Re-tune on
gpt-4.1-minionly if you are an existing customer and accept that its base will retire too
Most format and style fine-tunes are replaced by option 1 or 2 within a day, because a current model with a few dozen examples in a cached prompt usually beats a 2023 fine-tune of a much weaker model, and the cache makes the long prompt cheap. Option 3 matters for fine-tunes that encode real domain knowledge or a narrow skill that prompting cannot reach, and it is the only option that removes the dependency entirely. Our guide to the best open-weight models to self-host covers which models can be trained and served on your own hardware. Option 4 buys time, but it rebuilds on a platform OpenAI has said it is closing, so plan it as a bridge, not a destination.
gpt-image-1 moves to GPT Image 2.5, and the price model shifts
gpt-image-1 shuts down on October 23 and its listed replacements are GPT Image 2.5 Sunburst, which OpenAI describes as best "where editing precision matters most," and GPT Image 2.5 Flare, for "fast, high-quality everyday image generation" - OpenAI API changelog. Both use GPT Image 2 token rates: $8 per million image input tokens and $30 per million image output tokens, against $10 and $40 on gpt-image-1 - GPT Image 2.5 Flare model page. The per-token price falls by a quarter, but the number of tokens per image is a different matter: gpt-image-1 publishes per-image prices from $0.011 for a low-quality square to $0.25 for a high-quality portrait, while OpenAI notes that its GPT Image 2 calculator "does not estimate GPT Image 2.5 token consumption."
The new models also add xhigh and max quality settings above high. If your code passes quality="high", it will keep working; if it passes quality="auto", log the token usage of the first few hundred images, because the model now has two more expensive settings it could choose and the per-image cost is no longer published. The other three image models (gpt-image-1-mini, gpt-image-1.5 and chatgpt-image-latest) follow on December 1, so an image pipeline should move straight to 2.5 rather than to an intermediate model.
o1-pro becomes a mode, not a model
o1-pro was the most expensive model OpenAI served, at $150 input and $600 output per million tokens - OpenAI pricing. Its replacement is not a model ID: it is gpt-5.6-sol with reasoning.mode set to pro in the Responses API. OpenAI's reasoning guide explains that pro mode "performs more model work than standard mode" and bills that work at the selected model's standard token rates, and that GPT-5.6 and GPT-6 models support it. The per-token price drops by more than 95%, but token counts in pro mode are higher by design, so measure cost per completed task rather than comparing list prices.
Batch jobs, queued work and Azure deployments
Two operational cases catch teams that fixed their code in time. The first is queued work: a Batch API input file written on October 20 with "model": "gpt-4" in every line still names a dead model when it runs, and so does a scheduled job whose payload was serialized before the change. Search job queues and stored batch files the same way you searched code, and re-submit anything that will run after October 23. The second is hosting: if you call OpenAI models through Azure OpenAI, the dates on OpenAI's page do not apply. Microsoft runs its own lifecycle, with an 18-month standard window, an auto-upgrade for Standard deployments and none for Provisioned ones, and a 410 Gone response after retirement - Microsoft Foundry model retirements. Its example is gpt-4o version 2024-05-13, which retires on Azure on December 9, 2026 and auto-upgrades to gpt-5.6-sol on Standard deployments, seven weeks after the same model disappears from OpenAI's own API.
Why this matters: these four cases are where a "simple model swap" plan fails quietly, either by losing a fine-tune's behavior or by breaking after the deadline. How to apply it: give each special case its own row in the audit table, test fine-tune replacements against the fine-tune's original evaluation data, and treat queued and scheduled work as code that also needs migrating.
8. November 30 and December 11: Agent Builder, Evals, Prompts and GPT-5
The second wave of shutdowns is about products rather than models, and it lands on the teams that took OpenAI's managed tools furthest. On June 3, OpenAI announced that Agent Builder, the Evals platform and reusable prompt objects are all being deprecated, with shutdown on November 30 - OpenAI deprecations. ChatKit stays. These products were launched to make building agents easier without code, so the people using them are often the least prepared to rewrite them in code, which is what each migration path asks for.
The dates are close together but not identical. Existing evals become read-only on October 31, a week after the model shutdown, so any eval you want to re-run against a replacement model should be exported or recreated before then. Agent Builder, the Evals dashboard and API, and v1/prompts all stop on November 30. On the same day, Anthropic retires claude-sonnet-4-5-20250929 on the Claude API - Anthropic model deprecations, so a team that hedged by running both labs' mid-tier models from 2025 has two migrations due that week.
Agent Builder: export to code, then rebuild what does not transfer
OpenAI's migration guide offers two destinations: continue with the Agents SDK in your own application, or recreate the workflow as a ChatGPT Workspace Agent, which requires a Business, Enterprise or Edu workspace - Migrate from Agent Builder. Both start from the same export: open the workflow, select Code, choose Agents SDK, then TypeScript or Python.
The export in the screenshot is real code (a Zod schema and an Agent definition), which is the good news. The bad news is in the guide's own caveats: the process "does not convert your workflow graph or guarantee that every behavior transfers unchanged," and "workflows with strong determinism at their core may not migrate faithfully to a workspace agent." If your Agent Builder workflow was mostly branching logic with a model call at each node, the honest migration is to write the branching in ordinary code and keep the model calls, rather than handing a deterministic process to an agent that decides its own path. Our playbook on orchestrating parallel AI agents covers when an agent loop is the right shape and when a fixed pipeline is.
Evals: move to Promptfoo or your own harness
OpenAI recommends Promptfoo for evaluations and publishes a cookbook for the move. The move is manual: you recreate prompts, providers, test cases and assertions in a promptfooconfig.yaml, and the cookbook notes that its method "does not require an OpenAI Evals export feature" - OpenAI Cookbook, Evals to Promptfoo. Promptfoo is open source and runs locally or in CI, which is a real improvement for most teams, because evals that live next to the code get run on every change.
One fact belongs next to that recommendation: OpenAI agreed to buy Promptfoo in March 2026 - Bloomberg Law. The open-source project is still open source, and your config file is still yours, so this is not a reason to avoid it. It is a reason to keep your evaluation data (inputs, expected outputs, grading rules) in plain files in your own repository, so the harness can be swapped without losing the part that took months to build.
Prompt objects: move the text into your code
Reusable prompts let you reference a dashboard-managed template by ID and pass variables. The migration is to move the prompt text into your application and build the request yourself, which OpenAI frames as giving "more control over review, testing, deployment, and versioning" - Migrate from prompt objects.
# BEFORE: prompt managed in the OpenAI dashboard. Stops working Nov 30.
resp = client.responses.create(
prompt={"id": "pmpt_123", "version": "1",
"variables": {"customer_name": "Acme", "issue": "billing question"}},
)
# AFTER: prompt versioned in your repository, model chosen by your config.
from prompts import SUPPORT_TRIAGE_V1 # a plain string template in git
resp = client.responses.create(
model=settings.TRIAGE_MODEL,
instructions=SUPPORT_TRIAGE_V1,
input=f"Customer: {customer_name}\nIssue: {issue}",
)
Export the text of every prompt and every version you rely on now, while the dashboard still shows them. A prompt object also often pinned the model, so moving it into code is a natural moment to apply the model choice from section 5 instead of copying a retiring model ID out of the dashboard.
December 11: the original GPT-5 snapshots
The last date this year retires the GPT-5 snapshots from August 2025 and the o3 models: gpt-5-2025-08-07 moves to gpt-5.6-sol, gpt-5-mini-2025-08-07 to gpt-5.6-terra, gpt-5-nano-2025-08-07 to gpt-5.6-luna, and gpt-5-pro, o3 and o3-pro to gpt-5.6-sol (with pro mode for the pro models). Teams that migrated off GPT-4 to GPT-5 last year are in this group, which is the clearest evidence for the argument in section 12: moving to whatever is current is not a one-time fix, it is the first instance of a recurring task.
Why this matters: the November 30 products hold work that is harder to recreate than a model string: workflow logic, evaluation history and prompt versions. How to apply it: export Agent Builder workflows, eval definitions and prompt texts this week, put them in your repository, and schedule the rebuilds before November 30 rather than after the October 23 rush.
9. The Cost Math of Migrating
A forced migration is the one time every team re-prices its AI usage whether it planned to or not, and the outcome depends almost entirely on three choices: which model, which reasoning effort, and whether the stable part of the prompt is cached. List prices alone mislead in both directions. A move from gpt-4 to anything current looks like a large saving per token, and usually is, while a move from gpt-3.5-turbo to OpenAI's listed substitute looks like a modest price increase per token and can turn into a fifteen-fold bill once hidden reasoning is counted.
To make that concrete, the two examples below price typical workloads at the October 11 standard rates on OpenAI's pricing page - OpenAI pricing. The token counts are illustrative and stated in each example, and so are the reasoning-token assumptions, which you should replace with your own numbers from usage.output_tokens_details.reasoning_tokens after a test run. The point is the shape of the result, not the third digit.
Example 1: a classifier that ran on gpt-3.5-turbo
Take a pipeline that labels one million support tickets a month, sending about 400 input tokens and getting back a 5-token label. On gpt-3.5-turbo that costs $200 for input and $7.50 for output, or $207.50 a month. Moving it to the listed substitute, gpt-5.6-terra, at the default medium effort, and assuming each request spends 200 reasoning tokens deciding the label, costs $800 for input and $2,460 for output (205 tokens per request at $12 per million), or $3,260 a month. Setting reasoning_effort="none" on the same model brings it to $860. Moving to gpt-6-luna at none costs $42.50.
The spread between the most and least expensive options is about 80x, for a task where the cheaper models are more capable than gpt-3.5-turbo was. The Decisions API bar assumes the same 400 input tokens per request at $0.10 per million with no output charges; in practice the question definitions add some input, so treat it as a lower bound - OpenAI Decisions guide. The medium-effort bar is the one that ships by default if you copy the substitute from the deprecations table and change nothing else.
That default has a second-order effect that turns a cost problem into an outage. OpenAI simplified its usage tiers on October 6 to three: Build (after $5 of credit purchases) allows $500 of usage a month, Launch (after $100) allows $5,000, and Grow (after $500) allows $200,000 - OpenAI rate limits guide. A small team on the Build tier whose $207.50 classifier becomes a $3,260 classifier will hit the monthly limit in under five days, and since July OpenAI has also offered hard spend limits that return a 429 once reached. Either way, the migration that fixed the 404 produces a different failure within days.
Example 2: a chatbot that ran on gpt-4
Now take a support assistant that answers 100,000 questions a month with about 1,500 input tokens (a 1,000-token system prompt plus the conversation) and 300 output tokens. On gpt-4 that costs $4,500 for input and $1,800 for output, or $6,300 a month, which is why many teams moved off gpt-4 long ago and why the ones still on it are usually there by inertia rather than choice.
| Landing option | Assumptions | Input | Output | Monthly total |
|---|---|---|---|---|
gpt-4 (today) | no reasoning | $4,500 | $1,800 | $6,300 |
gpt-5.6-sol | effort none | $600 | $600 | $1,200 |
gpt-6.1-sol | effort low, 500 reasoning tokens assumed | $300 | $800 | $1,100 |
gpt-6.1-sol + caching | 1,000-token prefix cached at $0.10 per million | $110 | $800 | $910 |
gpt-6-luna | effort none | $15 | $15 | $30 |
Every row is a large saving, which is the real story for gpt-4 workloads: the deadline is an opportunity to cut the bill by 80% or more, not a cost. The spread inside the table matters too. Whether a support assistant needs a model at GPT-6.1 Sol's level or can run on GPT-6 Luna is a quality question that only your own evaluation set answers, and the price difference between those two choices is more than 30x. The cached row shows why prompt structure deserves attention during the migration: putting the stable system prompt first and the conversation after it cuts the input bill by almost two thirds on GPT-6.1 Sol, where cached reads cost 5% of the normal input price. For the full method of routing each request to the cheapest model that handles it, see our guide to cutting AI agent costs with model routing.
OpenAI's own DevDay session on cost, published on October 7, makes the same argument from the vendor's side: measure cost per completed task rather than per token, pick effort deliberately, and lean on caching and batch processing.
Promotional prices are not the price
Two of the prices in this guide have expiry dates, and both should be in your model of next quarter's bill. OpenAI's pricing page says GPT-5.6 Sol's current $4 / $20 price is promotional and "available at least through November 21, 2026"; OpenAI's August 21 changelog entry describes it as a 20% input and 33% output reduction, which implies $5 and $30 before the promotion. Google's pricing page shows Gemini 3.8 Flash at $0.75 / $3.75 through December 31 and $1.50 / $7.50 from January 1, 2027 - Gemini API pricing. A migration priced on promotional rates can look 20% to 50% cheaper than it will be in January.
Why this matters: the same deadline can cut one team's bill by 85% and multiply another's by fifteen, depending on choices that take minutes to make. How to apply it: price each workload on at least two landing models with effort set explicitly, use post-promotion prices, check the result against your usage tier's monthly limit, and set an alert on cost per request for the first week after the switch. Our guide to pricing your AI product to beat token costs covers how to pass these changes through to your own customers.
10. How Other Providers Retire Models, and What That Teaches
Every lab retires models, for the capacity reason in section 2, but they handle the moment of retirement in strikingly different ways. Some fail loudly, some forward your request to a newer model without asking, some publish a minimum lifetime up front. None of these is strictly better: each trades availability against predictability in a different place, and knowing the trade lets you choose a provider, or build your own policy, on purpose rather than by accident.
The comparison below uses each provider's own documentation as of October 11. It covers the policies that decide what your app experiences on the day a model goes away, which is the part that matters for planning.
| Provider | Minimum notice | What happens on the date | Published minimum lifetime |
|---|---|---|---|
| OpenAI API | 6 months GA, 3 months variants, ~2 weeks previews | Request fails (404 model_not_found) | None published |
| Anthropic Claude API | At least 60 days | Request fails | Yes: "not sooner than" date per model |
| Google Gemini API | Dates are "earliest possible" | Some retired IDs auto-route to the successor | No |
| Azure OpenAI | 18-month standard lifecycle | Standard deployments auto-upgrade, then 410 Gone | Yes: retirement date set at launch |
| OpenRouter | Varies by upstream | Depends on the upstream provider | expiration_date field when known |
Anthropic combines the shortest notice (at least 60 days) with the clearest floor: its deprecation table gives every active model a "not sooner than" date, such as September 28, 2027 for Claude Sonnet 5.5 and October 7, 2027 for Claude Haiku 5.5 - Anthropic model deprecations. It has also committed to preserving the weights of every publicly released model "for, at minimum, the lifetime of Anthropic as a company," which does not keep a model online but does keep the door open to bringing it back. Google lists shutdown dates as "the earliest possible dates on which a model might be retired," and for some models routes traffic forward automatically: requests to gemini-3.5-flash are sent to gemini-3.6-flash, and requests to gemini-3.7-flash to gemini-3.8-flash - Gemini API deprecations.
Azure is the most structured. Microsoft sets an 18-month retirement date when a model launches, publishes it through its Models API, auto-upgrades Standard deployments to the replacement and leaves Provisioned deployments to migrate by hand. For models it hosts from Anthropic, DeepSeek, Fireworks and Mistral, the standard lifecycle is 12 months - Microsoft Foundry model retirements. OpenRouter exposes an expiration_date on each model in its API, which is the right idea, but as of October 11 it is set on 18 models and on none of OpenAI's October 23 batch, so it would not have warned you about this one - OpenRouter models API. Vercel's AI Gateway still lists openai/gpt-3.5-turbo and openai/gpt-4-turbo in its model catalog on the same day - Vercel AI Gateway models. A gateway's catalog describes what it routes, not what the upstream will still serve.
Loud failure or silent substitution
The deepest difference in that table is what happens at the moment of shutdown. OpenAI and Anthropic fail the request, which is loud: your error rate spikes, someone gets paged, and nothing about your product's behavior changes without a human deciding it. Google (for some models) and Azure (for Standard deployments) substitute a newer model, which is quiet: your app stays up, but its output, latency and cost per request change on a date you did not pick, without a deploy, and possibly without anyone noticing. Neither is free. Loud failure costs availability; silent substitution costs control.
You do not have to accept your provider's choice. A gateway or a thin wrapper in your own code can implement either policy per call. Vercel's AI Gateway, for example, accepts a models array of fallbacks that it tries in order when the primary model fails, and bills the model that answers - Vercel AI Gateway model fallbacks.
// One request, an explicit fallback chain, through Vercel AI Gateway.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.AI_GATEWAY_API_KEY,
baseURL: "https://ai-gateway.vercel.sh/v1",
});
const response = await client.chat.completions.create({
model: "openai/gpt-6-luna",
messages: [{ role: "user", content: ticketText }],
// Gateway extension fields are not in the OpenAI SDK types.
...{
providerOptions: {
gateway: { models: ["openai/gpt-5.6-luna", "anthropic/claude-haiku-5.5"] },
},
},
});
A fallback chain like this turns a hard failure into a silent substitution, which is the right call for a support chatbot and the wrong one for a pipeline whose outputs feed billing or compliance. The rule that makes either safe is the same: log which model answered on every request, and alert when it is not the primary. A substitution you can see is a policy. A substitution you cannot see is the next incident.
Why this matters: the provider's retirement policy is part of the product you are buying, as much as price and quality. How to apply it: for each workload, decide whether it should fail loudly or fall back, implement that decision in your own code or gateway, and prefer providers that publish a minimum lifetime for anything you cannot afford to migrate on short notice. We made the broader case for keeping that control on your side in personal sovereign AI.
11. A Migration Runbook for the Next Twelve Days
On October 11 there are twelve days left, which is enough for a careful migration if the work starts now and runs in order. The sequence below is built around one constraint that disappears on October 23: while the old model still answers, you can run the same inputs through the old and new models side by side and compare. After the shutdown, the old model's behavior survives only in whatever outputs you saved, so the most valuable thing to do this week is capture it.
The plan assumes a small team with one or two OpenAI integrations. Larger organizations should run the same steps per integration in parallel, with one owner each, and keep a single shared audit table so nothing falls between teams. Every step names its output, because the outputs are what you will need if something breaks on the 23rd.
Days 1 and 2 (October 12-13): audit and capture
Run the three searches from section 3 across every repository, the Usage API query for the last 30 days, and a search of your request logs and job queues. Put the results in one table with a row per call site: model, location, owner, monthly requests, what the call does, and the planned landing model. Then pick 50 to 200 real, recent requests per call site (anonymized where needed) and save each one with the old model's actual output. That file is your evaluation set, and it is the only record of the old behavior that will exist after October 23.
Days 3 and 4 (October 14-15): choose and test the landing model
For each call site, choose a landing model with the decision path in section 5 and price it with the method in section 9. Then run the evaluation set through the new model with effort set explicitly and compare. Promptfoo can run the old and new model on the same tests in one command while both still exist, using its OpenAI providers (Chat Completions models are addressed as openai:chat:<model>) - Promptfoo OpenAI provider docs.
# promptfooconfig.yaml: same tests, old model vs two candidates
prompts:
- file://prompts/triage.txt
providers:
- id: openai:chat:gpt-3.5-turbo # baseline, gone Oct 23
- id: openai:chat:gpt-6-luna
config: { reasoning_effort: none, max_completion_tokens: 50 }
- id: openai:chat:gpt-5.6-terra
config: { reasoning_effort: none, max_completion_tokens: 50 }
tests: file://evals/triage_cases.yaml # inputs + expected labels
defaultTest:
assert:
- type: equals
value: "{{expected_label}}"
Look at three numbers per candidate: agreement with the expected outputs, cost per request (from the usage fields), and latency. If a candidate disagrees with the old model, read the disagreements before deciding anything: a meaningful share of them are usually cases where the old model was wrong.
Days 5 to 7 (October 16-18): make the code changes
Apply the changes in a fixed order so each one is testable on its own. Doing them all in one commit makes any regression impossible to attribute.
The order below runs from the change that touches the most code to the one that touches the most data, and each step should leave the app working against the old model, so you can merge them during the week without waiting for the switch. Only the parameter fixes depend on the landing model, and even those can live in the settings for each role, so the code that builds the request sends reasoning parameters only when the configured model takes them. That way the actual switch on day 8 is a one-value change to the setting from step 1, which is also your rollback.
- Centralize the model choice in one setting per workload, read from config
- Fix library defaults by passing the setting explicitly everywhere
- Fix parameters: effort,
max_completion_tokens, sampling, tool path A or B - Migrate stored values: database rows, saved agents, templates, queued jobs
- Add logging of the model that answered, reasoning tokens and cost per request
The fourth step is the one teams forget, because it is a data migration rather than a code change. If users or customers could ever pick a model, write a one-off script that maps every stored retiring ID to its replacement, run it against a copy first, and keep the mapping in code so any value written by an old client after the migration is translated on read. The fifth step is what makes the rest observable: without the model name and token counts in your logs, you cannot tell whether a cost spike on October 24 came from the new model, a fallback, or a prompt change.
Days 8 to 10 (October 19-21): roll out gradually
Ship the change behind a percentage rollout (5% of traffic, then 50%, then 100%) with at least a few hours at each step, and watch error rates, incomplete responses, cost per request and the evaluation agreement rate. Keep the old model as the rollback target for now, because it still works until the 23rd. Set alerts on HTTP 404 with model_not_found, on responses whose status is incomplete, and on cost per request above the level your pricing assumed.
The reason to roll out before the deadline rather than on it is that the old model remains a working rollback target only until October 23. A switch on October 19 that misbehaves can be reverted in minutes to a model that still answers; a switch forced on October 23 has nowhere to go back to except another new model. Use those days to look at real user-facing outputs too, not only dashboards: read a sample of fifty responses from each rolled-out workload, because tone and formatting changes (section 6) never show up as errors.
Day 11 (October 22) and after: sweep, then the next deadlines
On the last day, re-run the Usage API query for the previous 48 hours. Anything still calling a retiring model is a straggler: a script, a cron job, a teammate's notebook or a partner integration. Re-submit any batch jobs scheduled to run after the deadline. On the 23rd, watch the alerts. Then put the next dates on the calendar now, because they arrive faster than they look: export evals before October 31, rebuild Agent Builder workflows and move prompt objects before November 30, and move any original GPT-5 or o3 snapshot before December 11.
Why this matters: most migration failures come from order, not difficulty: changing everything at once, or losing the old outputs before comparing. How to apply it: copy this schedule into your tracker today with one owner per step, and treat the evaluation set from day 1 as a permanent asset, because you will run it again at every future retirement.
12. Design So the Next Retirement Is a Config Change
Everything above is the cost of one retirement. The chart in section 1 says there will be many more, so the highest-return work is not this migration but making the next one cheap. The goal is concrete: when OpenAI, Anthropic or Google announces the next shutdown, the change in your system should be one value in one config file, followed by your evaluation suite and a deploy. That is achievable with four pieces, none of which is exotic.
The underlying principle is the one that runs through this whole guide: a model is a supplier with a shutdown date, not a constant. Code that treats it as a constant (a string literal in forty files, a library default, a value baked into stored records) turns every supplier change into a code change. Code that treats it as a supplier puts it behind a seam, keeps a record of what it was supposed to do, and checks the supplier's schedule.
The four pieces
The first piece is a role registry: feature code asks for a role ("triage", "support-chat", "summarize") and one small module maps each role to a provider, a model and its settings (effort, output limit, tool path). The mapping lives in config, so changing the model behind "triage" touches nothing else. The second is an adapter layer that turns a role request into the right API call, so the difference between Chat Completions, Responses and another lab's API is handled in one place. The third is the evaluation suite you built in section 11, run in CI whenever the registry changes. The fourth is a CI check that fails the build if any config or code references a model on a known shutdown list.
# models.py: the only place model IDs live
import os
ROLES = {
"triage": {"provider": "openai", "model": os.environ.get("MODEL_TRIAGE", "gpt-6-luna"),
"effort": "none", "max_output": 50},
"support_chat": {"provider": "openai", "model": os.environ.get("MODEL_CHAT", "gpt-6.1-sol"),
"effort": "low", "max_output": 4000},
}
# Maintained from the providers' deprecation pages; CI fails on any match.
RETIRING = {
"2026-10-23": {"gpt-4", "gpt-4-turbo", "gpt-3.5-turbo", "gpt-4.1-nano",
"gpt-4o-2024-05-13", "o1", "o1-pro", "o3-mini", "o4-mini", "gpt-image-1"},
"2026-12-11": {"gpt-5-2025-08-07", "gpt-5-mini-2025-08-07", "gpt-5-nano-2025-08-07",
"gpt-5-pro-2025-10-06", "o3-2025-04-16", "o3-pro-2025-06-10"},
}
def check_no_retiring_models() -> None:
dead = set().union(*RETIRING.values())
bad = {role: cfg ["model"] for role, cfg in ROLES.items() if cfg ["model"] in dead}
if bad:
raise SystemExit(f"Roles point at retiring models: {bad}")
The defaults in that sketch are examples for the two workloads priced in section 9, not recommendations for yours; the right defaults are whatever your evaluation set says. The shutdown list should be maintained from the providers' deprecation pages. OpenAI publishes a Markdown copy of its page at the same URL with .md appended, which makes a weekly job that diffs it and opens a ticket a few lines of code. One more pattern is worth borrowing from Azure and Google without adopting their silence: let the registry accept a family alias ("openai-small-latest") that resolves to a concrete model in one table, and make every resolution visible in logs and, where users chose the model, in the product itself.
One concrete case is Founden, the platform publishing this guide. Its model policy file keys most choices by model family rather than by version, so a new release reaches every picker and default without a code change, and a stored choice of a retired model resolves to that family's newest release, with the user told which model actually ran. Its model catalog is regenerated from the providers' own model lists every six hours, and the last refresh before this guide was written ran at 06:07 UTC on October 11. The pattern is not specific to any platform: a family pointer plus visible substitution is a design you can build into the registry above in an afternoon.
Keep a second supplier warm
The last piece is the one most teams skip because it feels like overhead: keep at least one workload running on a second provider, or on open weights, all the time. The reason is not to switch in a crisis, which never goes as smoothly as planned, but to keep the adapter, the evaluation suite and your team's knowledge of the second API exercised, so a switch is a decision rather than a project. The scored table at the top of this guide gives Anthropic's models a 7 on control for exactly this reason: one more independent place to run your workload. An open-weight model such as gpt-oss-120b, which OpenAI releases under Apache 2.0 and describes as fitting on a single H100 - gpt-oss-120b model page, is the only option whose retirement date you set yourself.
Why this matters: the twelve-day scramble most teams are in now is the direct cost of having treated a model as a constant. How to apply it: build the role registry and the CI check as part of this migration rather than after it, keep the evaluation set in your repository, and schedule a monthly fifteen-minute review of every provider's deprecation page you depend on.
13. Conclusion: Where to Land
The October 23 shutdown looks like a model problem, and it is mostly a bookkeeping problem. The replacements are all far more capable than what they replace, most of them are cheaper per useful answer, and the code changes are well documented. What turns the deadline into an outage is not finding a call (a library default, a stored setting, a queued batch) or copying a substitute without setting its reasoning effort, which can multiply a bill more than tenfold and then hit a usage cap. The teams that handle it best are the ones that treat the next twelve days as the moment to build the seam, the evaluation set and the monitoring they will need for the 29 shutdowns OpenAI's page lists for this quarter.
For the decision itself, a short framework covers most cases. Read it as a set of starting points to test against your own evaluation set, not as a verdict: the right landing model is the cheapest one that passes your cases.
| If the retiring call... | Start with | Then |
|---|---|---|
| Classifies, routes or extracts | gpt-6-luna at effort none, or the Decisions API (beta) | Compare agreement with saved gpt-3.5-turbo or gpt-4 outputs |
Chats or drafts on old gpt-4 | gpt-6.1-sol and gpt-6-luna side by side | Keep the cheaper one that passes; cache the system prompt |
| Uses tools in Chat Completions and cannot be rewritten by Oct 23 | gpt-5.6-sol at effort none | Plan the Responses API move before the next deadline |
| Runs on a fine-tune | The replacement, prompted with cached examples | Retrain on a model you control if prompting falls short |
| Cannot be migrated on short notice, ever | A provider with a published minimum lifetime, or open weights | Keep a second supplier warm |
The table leans toward the newer, cheaper models because the numbers in sections 5 and 9 do: for this batch, the newest small model beats the oldest large one on capability and costs a fraction of it. The one place it leans toward OpenAI's listed substitute is tool calling in Chat Completions, because that is where a rewrite under deadline is the bigger risk.
None of those choices is permanent, and that is the point. Every model in that list will itself be retired, most of them within two years if the current pace holds, and the GPT-5 snapshots retiring on December 11 were the "new model" that teams migrated to last year. A migration done well leaves behind a registry, an evaluation set and a habit of reading deprecation pages, which makes the next one a config change. A migration done in a hurry leaves behind forty new string literals and a fresh shutdown date. If you want a wider view of which models are worth building on right now, our guides to GPT-6 Astra's real pricing, what shipped with GPT-6 Astra and Claude Haiku 5.5 against GPT-6 Luna go deeper on each candidate, and our security checklist for AI-built apps covers the other things worth checking while the code is open.
This guide reflects OpenAI's deprecations page, model pages and pricing, plus the Anthropic, Google, Microsoft, OpenRouter and Vercel documentation cited, as of October 11, 2026. Shutdown dates, substitutes, promotional prices and library defaults change frequently, so verify current details on each provider's page before migrating production traffic.