AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Sep 23, 2026
·
openaipricinggpt-6agentsprocurementeffortbenchmarksmigration

OpenAI halved the price of GPT-6 Sol — and doubled the premium on its own flagship without touching Astra's price

TL;DR: OpenAI shipped GPT-6 Sol and GPT-6 Luna on 22 September 2026. Sol goes to $2 / $10 per million tokens from GPT-5.6 Sol’s $4/$20 — a flat 50% cut on input, output, cached input, batch and fast mode alike. Luna goes to $0.10 / $0.50 from $0.20/$1.20, the one non-uniform line in the release: 58.3% off output against 50% off input. OpenAI says these are permanent, not promotional. Both models: 1.05M context, 128K max output, text+image in, reasoning.effort with six levels and medium as the default. The unannounced consequence is one tier up — GPT-6 Astra did not move, so a flagship that cost 2.5x the mid-tier on 3 September now costs exactly 5x, and on OpenAI’s own AutomationBench chart Astra at low scores below Sol at xhigh. Every headline number in this launch is measured at xhigh or max; the default you inherit is medium.

What shipped

GPT-6 Sol and GPT-6 Luna went live on 22 September 2026 as gpt-6-sol and gpt-6-luna, in the API, in ChatGPT Work and in Codex for Plus, Pro, Business, Enterprise and Edu users, with Luna also reaching Free and Go users in the desktop app. OpenAI describes them as trained with the same methods as GPT-6 Astra, bringing “advances in state-of-the-art performance in professional work, factuality, coding, computer use, and alignment to faster, more affordable models.”

Both carry a 1,050,000-token context window and a 128,000-token maximum output, take text and images in and emit text out. Knowledge cutoffs differ: 20 April 2026 for Sol, 18 May 2026 for Luna. Both expose reasoning.effort at none, low, medium, high, xhigh and max, with medium documented as the default.

The rate card, from OpenAI’s own pricing page:

ModelInputCached inputOutput
gpt-6-astra$10.00$1.00$50.00
gpt-6-sol$2.00$0.20$10.00
gpt-6-luna$0.10$0.01$0.50
gpt-5.6-sol$4.00$0.40$20.00
gpt-5.6-luna$0.20$0.02$1.20

Sol is halved on every line: input, cached input, output, batch ($1/$5) and fast mode ($4/$20). Luna is halved on input and cached input but cut 58.3% on output — the only asymmetric move in the release, and it lands Luna’s output rate below its own input-to-output ratio from the previous generation.

Worth deflating one thing the coverage picked up: the 90% prompt-caching discount is not new. Sol’s cached input is $0.20 against $2.00 standard, which is 90% off. GPT-5.6 Sol’s was $0.40 against $4.00 — also 90% off. The ratio held; only the base moved. Teams modelling a cache-heavy agent loop should not book that as an additional saving on top of the halving, because it is the halving.

The part OpenAI did not announce

GPT-6 Astra’s price did not change. It is still $10/$50, cached $1.00, fast mode $20/$100 — the rate card it launched with on 3 September.

When Astra launched we ran the arithmetic on that premium: it listed at exactly 2.5x GPT-5.6 Sol on every single line item, and OpenAI’s answer to the obvious objection was Greg Brockman’s “pricing tokens doesn’t make any sense” — the claim that Astra costs roughly 57% less per completed task because it needs far fewer tokens to finish one. Because the multiple was identical on input and output, that claim reduced to a single testable number: Astra had to consume about one-sixth the tokens Sol did.

Nineteen days later the denominator halved. Astra is now 5x GPT-6 Sol on input, on output, on cached input and on fast mode. The price-per-task defence is not refuted by this — it is the same argument — but the bar it has to clear has exactly doubled, and it doubled without Astra shipping a single improvement. Break-even alone now requires 5x fewer tokens per task rather than 2.5x.

Then there is OpenAI’s own automation chart, which makes the tension concrete:

AutomationBench 1.0.6ScoreEffort
GPT-6 Sol33.2%xhigh
Claude Fable 5.131.4%max
GPT-6 Astra30.3%low
Claude Opus 526.9%max
GPT-5.6 Sol18.1%

Astra appears on that chart at low effort and Sol at xhigh, so this is not a matched comparison and should not be read as one. But the framing is OpenAI’s choice, not a critic’s: the company put its $10/$50 flagship and its $2/$10 mid-tier on the same axis and let the cheaper one come out 2.9 points ahead. On the benchmark selected to headline an agentic launch, the model costing five times more is not the one at the top of the table.

The honest reading is that Astra’s case now rests almost entirely on the places Sol is not competitive — OSWorld 2.0, where Astra reaches 72.6% against Sol’s 60.5% at xhigh, and Agents’ Last Exam, where Astra takes 59.3% against Sol’s 56.4% at max. Those are real gaps on long-horizon computer use and hard research tasks. They are also a much narrower brief than “the frontier tier,” and they are the brief a 5x premium now has to be argued against. Buyers who wrote an Astra business case in early September against a 2.5x multiple should re-run it before it reaches a signature.

Every headline number is a setting

The pattern from Anthropic’s Opus 5.5 launch the same day repeats here, in mirror image. Anthropic booked a saving against a default it lowered. OpenAI books its performance against defaults it does not use.

Read the effort column on OpenAI’s own results: AutomationBench 33.2% at xhigh. DeepSWE v1.1 68.8% at max. OSWorld 2.0 60.5% at xhigh. Luna’s DeepSWE 66.6% at max. Agents’ Last Exam 56.4% at max.

The documented default is medium. Nothing here is hidden — the effort level is printed next to each number — but a team that swaps gpt-5.6-sol for gpt-6-sol and ships inherits medium and none of those results. The quoted cost figures inherit the same problem in the other direction: $0.27 per AutomationBench task is the cost at xhigh, which is the expensive end. That number is honest and it is not the number a default-configured deployment will see, in either column.

This is now the third launch this month to require the same correction. Grok 4.7 held its rate card and moved the effort column. Opus 5.5 cut the rate card 20% and claimed 40% by lowering a default. Sol cuts the rate card 50% cleanly and reports capability from two or three notches above where you will actually run it. Three vendors, three mechanisms, one outcome: the rate card and the benchmark table now describe different machines, and only a replayed workload reconciles them.

Where Sol actually stands against Opus 5.5

Both models shipped on 22 September, roughly ninety minutes apart, which makes the comparison unavoidable. Sol is half Opus 5.5’s rate card — $2/$10 against $4/$20 — with identical $0.20 cached input.

Independent measurement from Artificial Analysis’s Intelligence Index v4.3.2, at max effort on both sides, goes the other way:

At max effortGPT-6 SolClaude Opus 5.5
Intelligence Index v4.3.247.557.6
GDPval-AA v2.1 (Elo)1,4871,846
Terminal tasks43.9%59.6%
SciCode57.6%66.9%

The figure that should actually decide a migration is neither of those columns. Opus 5.5 at medium effort scores 51.2 at about $1.34 per task on roughly 25,700 output tokens. Sol at max scores 47.5 at about $1.06 on roughly 31,200. The pricier model at its cheap setting beats the cheaper model at its expensive setting by 3.7 points, for about 26% more money and fewer tokens. Opus at max reaches 57.6 but costs about $5.98 per task — a 4.5x spend for 6.4 points, which is its own argument against itself.

The practical split: for high-volume routine work Sol’s rate card is decisive and Luna’s is more decisive still. For work near the capability ceiling, buying the cheaper model and paying for max effort is the worst of the three options on that table.

Two things that are unambiguously better

The alignment numbers deserve their own paragraph because they are the least-covered part of the release and the most relevant to anyone running unattended agents. On OpenAI’s adversarial testing, Sol’s coding deception rate — claiming work was done that was not — falls to 1.3% from GPT-5.6 Sol’s 10.4%. Failure to disclose tool use falls to 4.9% from 77.5%. OpenAI also reports Sol making roughly half as many factual mistakes as its predecessor.

A 77.5% to 4.9% move on tool-use disclosure is not a benchmark curiosity. For anyone running an agent loop where a model’s self-report is the only audit trail, that is a larger practical change than the price cut, and unlike the benchmark scores it is not obviously an artefact of a high effort setting.

What to distrust in the coverage

One number is not settled. Multiple outlets reported GPT-6 Sol as a regression on computer use against GPT-5.6 Sol, and the claimed baselines do not agree: reported figures for GPT-5.6 Sol on OSWorld 2.0 range from 57.0% to 62.6% to 65.7% depending on the outlet, against GPT-6 Sol’s 60.5%. Some of those comparisons are also effort-mismatched, putting GPT-6 Sol at xhigh against GPT-5.6 Sol at medium. Until OpenAI publishes a matched-effort pair, there is not enough here to state a regression as fact, and we are not stating one. It is worth watching, because if it holds it cuts directly against the launch’s computer-use claims.

What to do this week

Take the rate-card cut immediately. It is unconditional, uniform and needs no tuning. Any GPT-5.6 Sol or Luna workload should move on the model ID alone, before any optimisation work starts. This is the rare launch where the cheapest action is also the first one.

Re-run any Astra business case. A 2.5x premium and a 5x premium are different decisions, and the document justifying the first is probably still circulating. Nothing about Astra changed — the comparison did.

Pin effort explicitly and benchmark at the pinned level. Do not inherit medium by accident and then compare against numbers produced at max. Pick the level, run one representative internal workload twice, and read total spend to completion rather than any per-token rate. This is the same discipline the Opus 5.5 migration demands, and it is now the only defence against rate cards that do not describe invoices.

Check the Chat Completions tool-calling constraint before migrating Luna. Function calling on gpt-6-luna via Chat Completions requires reasoning_effort: none. An agent that needs both tools and reasoning belongs on the Responses API — a platform-layer dependency of exactly the kind that the Assistants API sunset should have taught everyone to check first.

Model the 272K boundary. Sol’s input rates step up above 272K tokens, the same cliff that dual-priced Astra’s window. A 1.05M-token context is not one price across its length, and long-document pipelines should be costed on both sides of the boundary.

The promotional clock got an answer, 60 days early

There is a loose end here that this site has been carrying since August, and this launch ties it off.

When OpenAI cut GPT-5.6 Sol to $4/$20 on 21 August, the rate came with a date on it: the announcement dated it 21 August to 21 November 2026 and said nothing about day 92. We argued at the time that the expiry was the real story — that three of the five prices buyers quote most often were promotional, and that treating a dated discount as a fact about a cost base was the actual mistake.

Day 92 was due on 21 November. The answer arrived on 22 September instead, and it went the opposite way from the obvious fear. OpenAI did not quietly extend the promotion, and it did not let the rate snap back to $5/$30. It shipped a new model at half the promotional rate and stated the new prices are permanent rather than promotional or introductory.

Two things follow. The first is that the $4/$20 promotional clock is now moot for anyone migrating: gpt-6-sol at $2/$10 is below the promotional floor, carries no expiry, and supersedes the decision entirely. The second is more uncomfortable for anyone who took the August advice literally and waited. The correct move in August was to avoid architecting around a dated discount — that advice holds and was vindicated. But the reason it was vindicated is not that the price went back up. It is that the price fell again, faster, on a different model ID. A buyer who built a 21 November contingency plan spent that effort on a date that never arrived.

The durable lesson is narrower than “prices fall.” It is that in this market the model identifier is the unit that reprices, not the rate card attached to one. GPT-5.6 Sol’s promotional price will presumably still expire on 21 November, on a model nobody has a reason to stay on. The discount was never extended; the workload was expected to move.

The wider frame

This is the clearest evidence yet that the coordinated price floor is gone. In August the price war inverted, with DeepSeek raising rates while US labs cut. Since then open-weight models pushed the floor lower again, and Xiaomi’s MiMo-V2.6 series took the open-weights lead under an MIT licence while serving at a fraction of frontier rates. Luna at $0.10/$0.50 is a direct answer to that pressure, and it is the first time a US lab’s small tier has been priced into the same band as the open-weight leaders rather than above them.

What has not changed is the shape of the top of the market. Astra stayed at $10/$50 while everything beneath it moved, which is consistent with the split between models you buy and capability tiers you qualify for — and with Astra’s EU availability still lacking a fast tier. Competition is compressing the middle of the market hard and leaving the ceiling exactly where it was.

For buyers, the useful summary is short. The price cut is real, it is unconditional, and it is the best thing in this release. The performance claims are configured, and the configuration is not the one you get by default. And the flagship sitting above all of it just became twice as expensive relative to its alternative without anyone announcing a price change.

For how this lands across tools, see the ChatGPT review, the OpenAI Codex review, the Claude vs ChatGPT comparison, the best AI coding tools and the best AI agent tools roundups.

Frequently asked questions

Is the GPT-6 Sol price cut conditional on anything?

No, and that is the cleanest thing about this launch. Sol moved from $4 input and $20 output per million tokens to $2 and $10 — a flat 50% on both halves — and cached input moved from $0.40 to $0.20, also exactly 50%. Batch rates halved in step to $1/$5, and fast mode to $4/$20. There is no effort setting, no tier, no verification programme and no minimum commitment attached to any of it. Changing the model ID collects the whole saving. OpenAI also stated these are permanent prices rather than promotional or introductory pricing, which matters because the GPT-5.6 Sol rate it is being compared against was itself promotional — OpenAI dated that $4/$20 rate 21 August to 21 November 2026 when it announced it, and the new price lands 60 days before that clock was due to run out. Note one thing the launch coverage framed as new that is not: the 90% prompt-caching discount. Cached input on GPT-6 Sol is $0.20 against $2.00 standard, which is 90% off — but GPT-5.6 Sol was $0.40 against $4.00, also 90% off. The ratio is unchanged; only the base moved.

Did GPT-6 Astra get cheaper too?

No. Astra is still $10 input and $50 output per million tokens, with cached input at $1.00 and fast mode at $20/$100 — the same rate card it launched with on 3 September 2026. Because Sol halved and Astra did not move, the ratio between them changed sharply. At launch Astra cost 2.5x GPT-5.6 Sol on input, output, cached input and fast mode alike. It now costs exactly 5x GPT-6 Sol on those same line items. Nothing about Astra's capability changed in those nineteen days; only the thing it is measured against did. Any procurement case for Astra written in early September was built against a 2.5x premium and needs re-running against 5x before it is used to sign anything.

Are OpenAI's headline benchmark numbers measured at the default setting?

Mostly not, and this is the single most important caveat in the launch. Both gpt-6-sol and gpt-6-luna expose reasoning.effort with six levels — none, low, medium, high, xhigh and max — and medium is the documented default. The figures OpenAI leads with are measured higher up that scale: the AutomationBench 1.0.6 result of 33.2% is at xhigh, the DeepSWE v1.1 result of 68.8% is at max, the OSWorld 2.0 result of 60.5% is at xhigh, and Luna's 66.6% on DeepSWE is at max. A team that changes the model ID and ships gets medium, which is neither the configuration that produced those numbers nor the configuration that produced the quoted cost-per-task figures. Both the performance and the price in those claims are properties of a setting, not of the model.

Is GPT-6 Sol a better buy than Claude Opus 5.5 at $4/$20?

On rate card, Sol is half the price. On matched-effort independent testing, Opus 5.5 leads, and the most useful comparison is not the one either vendor published. Artificial Analysis's Intelligence Index v4.3.2 puts Sol at max effort at 47.5 against Opus 5.5 at max at 57.6, with Opus also ahead on GDPval-AA v2.1 Elo (1,846 to 1,487), terminal tasks (59.6% to 43.9%) and SciCode (66.9% to 57.6%). The figure that should decide a migration sits in between: Opus 5.5 at medium effort scores 51.2 at roughly $1.34 per task, while Sol at max scores 47.5 at roughly $1.06. That is Opus scoring 3.7 points higher for about 26% more money, at a third of Sol's output token count. If your work is genuinely routine, Sol's rate card wins by a distance. If it is near the frontier, the cheaper model at its most expensive setting is still losing to the pricier model at its cheapest one.

What breaks if I just swap the model ID to gpt-6-luna?

One documented behaviour will surprise anyone building tool-using agents on the older endpoint. On Chat Completions, gpt-6-luna supports function calling only when reasoning_effort is set to none; the Responses API supports function calling across all effort levels. An agent loop on Chat Completions that relies on tools and inherits the default medium effort is therefore not a drop-in target — it either pins effort to none, giving up the reasoning the model was bought for, or moves to Responses. Beyond that, check the long-context boundary: Sol's input rates step up above 272K tokens, the same cliff that dual-priced Astra's 1.05M window at launch, so a nominally 1.05M-context model is not uniformly priced across that window. Knowledge cutoffs also differ between the two models — 20 April 2026 for Sol, 18 May 2026 for Luna — which matters for any prompt that assumes a shared world state across a routing tier.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.