OpenAI fixed a vision bug in GPT-6 Sol and Luna three days after launch — and neither model has a dated snapshot to tell the two versions apart
TL;DR: OpenAI released GPT-6 Sol and GPT-6 Luna on 22 September 2026. On 25 September its API changelog recorded: “Fixed a bug in image encoding that degraded image understanding in GPT-6 Sol and GPT-6 Luna,” noting the fix improves visual tasks in the API and Codex, including computer use, and advising developers with image inputs to rerun evaluations. The disclosure is clear and creditable. The problem is that neither model has a dated snapshot — OpenAI’s reference pages list exactly one identifier each, gpt-6-sol and gpt-6-luna. So the degraded build and the fixed build answer to the same string. You cannot pin the old one, reproduce it, or roll back to it. Every vision and computer-use number published in that 72-hour window — including, potentially, OpenAI’s own OSWorld 2.0 launch figure — measured a defect. Record a timestamp next to every model ID, and treat a missing snapshot alias as a procurement finding.
A three-line changelog entry with a large blast radius
OpenAI’s API changelog for 25 September 2026 says this:
“Fixed a bug in image encoding that degraded image understanding in GPT-6 Sol and GPT-6 Luna. This update improves results on visual tasks in the API and Codex, including computer use.”
The entry goes on to recommend that developers with image input use cases rerun their evaluations and retry affected workflows.
Start with the credit, because it is deserved. OpenAI could have shipped this quietly. Encoding fixes are the kind of change that routinely lands with no note at all, and a vendor that writes down “our model was worse than it should have been for three days, go re-measure” is behaving better than the industry baseline. The entry is specific about the defect, specific about which models, specific about which workloads, and explicit about what developers should do.
Now the timeline. Sol and Luna were released on 22 September, at $2/$10 and $0.10/$0.50 per million tokens — a flat 50% cut against the GPT-5.6 tier. A halved price on a frontier-class model is the single most reliable trigger for evaluation work in this industry. Within 72 hours, benchmark write-ups, vendor bake-offs, migration spikes and “is the cheap one good enough now” analyses were all in flight simultaneously.
That 72-hour window is precisely the interval in which the models’ image understanding was degraded.
The part with no workaround
If this were only a bug and a fix, it would be unremarkable. Software has defects; vendors fix them. What makes it structural is what OpenAI’s own model reference pages show.
The page for gpt-6-sol lists one snapshot identifier: gpt-6-sol. It instructs developers to “use gpt-6-sol in your API requests”. There is no gpt-6-sol-2026-09-22. The page for gpt-6-luna is identical in this respect — one string, no dated aliases. Both models carry a 1,050,000-token context window and accept text and image input.
So the model that degraded your image understanding on 23 September and the model that does not degrade it on 26 September are the same string. Not similar strings. The same one.
Three capabilities that mature engineering practice assumes are therefore unavailable:
- You cannot hold a version steady. The standard way to qualify a new model is to pin the current one, test the candidate, then switch deliberately. With a floating-only identifier, the thing you are holding steady can move underneath you mid-qualification.
- You cannot reproduce a prior result. A benchmark number from 23 September describes a build that no longer exists and has no name. Nobody — not you, not OpenAI, not an independent evaluator — can re-run it.
- You cannot roll back. This particular change was an improvement on average, which makes rollback sound irrelevant. It is not. A fix to image encoding changes how images become tokens, and prompts tuned against the old encoding can regress even as aggregate quality rises. Teams that spent three days adjusting prompts to compensate for degraded vision have just had their compensation invalidated, in the opposite direction, with no way to revert while they re-tune.
This is a genuine departure from the convention OpenAI itself established. The GPT-3.5 and GPT-4 generations shipped dated snapshots as a matter of course — gpt-3.5-turbo-1106, gpt-4.1-nano-2025-04-14 — with the undated alias offered as a convenience for people who wanted the latest. Those dated strings are load-bearing infrastructure: they are what makes a deprecation calendar expressible at all, which is why this week’s shutdown list is written in snapshot IDs rather than in model families. The GPT-6 tier has dropped them.
The question OpenAI’s launch numbers now carry
OpenAI’s launch materials led with agentic and computer-use results. Among them, per this desk’s record of the launch, an OSWorld 2.0 score of 60.5% at xhigh effort and an AutomationBench 1.0.6 score of 33.2% at xhigh.
OSWorld is a desktop-automation benchmark. The model is shown screenshots and asked to operate a computer. Its entire perception of the task is images — exactly the pathway the encoding bug degraded, and exactly the workload the changelog names when it says the fix “improves results on visual tasks in the API and Codex, including computer use”.
There are only two possibilities, and OpenAI’s changelog does not say which holds:
- The launch figures were measured on the shipped, degraded build. In that case Sol’s real computer-use capability is better than OpenAI advertised, and the published numbers understate the product.
- The launch figures were measured on an internal build that already had the fix. In that case the numbers were true of OpenAI’s model but not of the endpoint developers could actually call between 22 and 25 September.
Neither is scandalous. Both matter to anyone who used those figures to decide something, because in case 1 the benchmark is stale-low and in case 2 the benchmark never described the shipped artefact. And the asymmetry is worth noticing: a vendor bug that makes a vendor’s own launch numbers look worse than reality is the rare error that argues for the vendor’s honesty rather than against it.
Text-only results are not implicated. An image-encoding defect does not touch DeepSWE, SciCode or terminal-task scores, so those can be read at face value. This is also the clearest available argument for why independently measured index scores on private test sets are worth more than launch-day slides: an outside evaluator re-running on a schedule catches a step change like this one. A launch deck, by construction, describes a single moment that may already have passed.
Who was affected without knowing it
The teams most exposed here are the ones that do not think of themselves as doing vision work. Image inputs enter production pipelines through paths nobody labels:
- A document step that rasterises PDF pages before reading them.
- A support workflow that accepts customer screenshots.
- An agent that screenshots a browser to decide its next click — the core loop of computer-use tooling.
- A coding assistant reading a design mock, a failing chart or a screenshot of a stack trace. The changelog names Codex specifically, which means Codex users were in scope whether or not they considered themselves to be sending images.
For developers auditing this, the reliable query is against request logs rather than source code: filter the 22–25 September window for requests carrying image content parts. Searching the codebase for the word “vision” will miss most of them. Anything that comes back needs re-running — and then a second check, on whether a decision has already been made on those numbers. A routing rule, a model selection, a migration sign-off or a published comparison built on pre-fix measurements is now resting on a defect.
There was also independent friction around Sol and Luna’s image handling in this period on third-party surfaces, including reports of image inputs being rejected outright for these models on Bedrock. Those are integration bugs in other people’s adapters rather than OpenAI’s encoding defect, but they compound the same conclusion: the image path on a three-day-old model tier is the least settled part of it.
The thesis, stated plainly
A model ID is not a version number, and this week it stopped even pretending to be one.
OpenAI did the disclosure well and the versioning badly, and only the second half is durable. The bug is fixed; the fix is an improvement; anyone who reruns their evaluations today gets an honest picture of a better model. But the condition that produced the problem is still in place: a frontier tier with image input, sold on computer-use performance, exposing exactly one unversioned string per model. The next silent change behind gpt-6-sol will be just as unpinnable, and there is no guarantee the next one is an improvement or that it comes with a changelog entry.
Capability moves under stable names in both directions. It moved up here. It has moved down elsewhere this month, by routing rather than by bug, and the same month saw image models ship new quality settings behind unchanged token prices. In none of those cases does the rate card or the version string tell you anything happened. The premium-tier procurement arguments that dominate ChatGPT-platform buying decisions are all built on benchmark deltas between named models — and a delta between two numbers is meaningless if one of the names quietly referred to something else when it was measured.
The cheap fix is a timestamp. Write the UTC time next to the model ID in every evaluation artefact you keep, treat that pair as the identity of what you tested, and run a small regression suite on a schedule instead of only at migration time. A dozen fixed prompts with known-good outputs, including image cases, would have caught this on 25 September without anyone reading a changelog at all.
Frequently asked questions
What exactly did OpenAI fix on 25 September 2026?
OpenAI's API changelog entry for 25 September 2026 reads: 'Fixed a bug in image encoding that degraded image understanding in GPT-6 Sol and GPT-6 Luna. This update improves results on visual tasks in the API and Codex, including computer use.' The entry recommends that developers with image input use cases rerun their evaluations and retry affected workflows. Three things in that wording matter for a buyer. First, the defect was in image encoding — the stage that turns your image into tokens the model can read — which means it affected every image-bearing request rather than some subset of difficult ones. Second, the scope explicitly includes Codex and computer use, which are screenshot-driven workloads where the model's only view of the world is an image. Third, OpenAI itself says to rerun evaluations, which is a vendor telling you in plain language that results you collected before that date do not describe the model you are now calling. The disclosure is honest and specific, and OpenAI deserves credit for publishing it rather than shipping the fix silently. The gap is not in the disclosure. It is in the absence of any identifier that lets you act on it precisely.
Which builds of GPT-6 Sol can I pin, and can I go back to the old one?
You cannot pin either build, and you cannot go back. OpenAI's model reference page for gpt-6-sol lists exactly one snapshot identifier — gpt-6-sol — and instructs developers to 'use gpt-6-sol in your API requests'. There is no gpt-6-sol-2026-09-22, no dated alias of any kind, and the same is true of gpt-6-luna. This is a real departure from the convention OpenAI itself established across the GPT-3.5 and GPT-4 generations, where dated snapshots like gpt-3.5-turbo-1106 and gpt-4.1-nano-2025-04-14 were the norm and the floating alias was the convenience option. Those dated strings are the reason a deprecation calendar can even exist, and it is why the current shutdown list is written in snapshot IDs rather than model families. Without them, three specific capabilities disappear: you cannot hold a model steady while you qualify a replacement, you cannot reproduce a benchmark result from a fortnight ago, and you cannot roll back if a fix changes behaviour your prompts depended on. A fix is still a behaviour change, and a behaviour change you cannot opt out of or revert is an uncontrolled dependency no matter how much it improves average quality.
Are OpenAI's own launch benchmarks affected?
At least the vision-dependent ones raise a question OpenAI has not answered, and the two possible answers are both worth knowing. The launch materials on 22 September led with computer-use and automation results, including an OSWorld 2.0 figure of 60.5% at xhigh effort and an AutomationBench 1.0.6 figure of 33.2% at xhigh, as recorded in this desk's coverage of the launch pricing. OSWorld is a screenshot-driven desktop benchmark; the model sees images and nothing else. If those figures were measured on the build that shipped on 22 September, they were measured through the encoding defect, and the model's real capability is higher than OpenAI advertised — an unusual direction for a vendor to err in, and one that argues the published numbers understate the product. If instead they were measured on an internal build that already had the fix, then the numbers were accurate about OpenAI's model but not about the artefact developers could call for those three days. There is no third option, and OpenAI's changelog does not say which applies. Text-only results — DeepSWE, SciCode, terminal tasks — are not implicated by an image-encoding bug and can be read normally. This distinction, between a benchmark that describes a vendor's build and one that describes the endpoint you pay for, is the same gap that makes independent index scores worth more than launch-day slides.
Do I need to rerun my own evaluations, and how do I know if I was affected?
If any evaluation you ran between 22 and 25 September 2026 sent an image to gpt-6-sol or gpt-6-luna, treat the result as void and rerun it. OpenAI's own guidance says exactly this. The harder problem is that many teams will not know whether they were affected, because image inputs enter modern pipelines through paths nobody labels as vision work: a document-understanding step that rasterises PDF pages, a support workflow that accepts customer screenshots, an agent that takes a screenshot to decide its next click, a coding assistant reading a design mock or a failing chart. Any of those is an image request. The practical audit is to query your request logs for the window and filter for requests carrying image content parts rather than searching your code for the word 'vision'. Anything that comes back needs rerunning; anything text-only does not. Then go one step further and check whether the conclusion you drew from those evaluations has already been acted on — a model selection, a routing rule, a migration sign-off or a published comparison made on pre-fix numbers is now resting on measurements of a defect.
What should change in how a team evaluates models after this?
Three changes, none of them expensive. First, record a UTC timestamp alongside the model ID in every evaluation artefact you keep, and treat the pair as the identity of what you tested. A bare model name is no longer a sufficient description of a measurement when the vendor ships unversioned changes behind it. Second, prefer dated snapshots wherever a vendor offers them, and treat their absence as a procurement finding rather than a detail — if a model family has no pinnable snapshot, you have accepted that your supplier can change your production behaviour without an identifier changing, and that belongs in a risk register. Third, run a small standing regression suite against the models you depend on, on a schedule rather than only at migration time, and include image cases if you send images. A dozen fixed prompts with known-good outputs, run daily, would have surfaced a step change in vision quality on 25 September without anyone reading a changelog. The broader lesson generalises past this incident: capability can move under a stable model name in either direction, whether through a fix like this one or through routing changes that silently downgrade what you get at peak hours. Buyers who monitor only price and version strings are watching the two variables that happen to be easiest to see.
Sources
- OpenAI — API changelog (primary: the 25 September image-encoding fix, verbatim; the 22 September Sol and Luna release and pricing)
- OpenAI — gpt-6-sol model reference (primary: single snapshot identifier, context window, image input)
- OpenAI — gpt-6-luna model reference (primary: single snapshot identifier, context window, image input)
- OpenAI — Introducing GPT-6 Sol and Luna
- OpenAI — API pricing
- TechCrunch — OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes
- Vellum — GPT-6 Sol and Luna benchmarks explained
- crewAI issue #7732 — image input_files rejected for GPT-6 Sol/Luna on Bedrock
- OpenAI Developer Community — Announcing GPT-6 Sol and GPT-6 Luna in the API, Codex and ChatGPT
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.