AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Oct 5, 2026
·
openaianthropicpricinggpt-6-1-solclaudeapibenchmarksprocurementcost-modellingdevelopers

GPT-6.1 Sol and Claude Sonnet 5.5 both cost $2/$10, land a day apart, and bill an order of magnitude differently

Correction (5 October 2026): an earlier version said OpenAI advertised no batch discount for GPT-6.1 Sol. OpenAI’s pricing page lists Sol’s batch and flex rates at $1 input / $5 output, 50% off standard, the same discount Anthropic offers. Batch is therefore not a difference between the two; the cache-read rate and the 272K long-prompt surcharge still are. Corrections log.

Update (1 October 2026) — the other two things DevDay did to the bill. This article covers the model prices OpenAI set on 29 September. Two further changes from the same keynote move money for people who never touch the API. First, the subscription ladder was restructured: Pro 500 arrived at $500/month with 25x the Plus allowance and sole access to the Ultrafast tier, while Pro 200 reopened with its included usage cut from 20x to 10x effective 30 October 2026 — and once you divide price by allowance, every tier now costs $20 per 1x of Plus, with the old Pro 200 at $10 being the only volume discount in the lineup and the one discontinued. Second, OpenAI launched dots, always-on agents on GPT-6 Astra, and Pro access excludes the European Economic Area, Switzerland and the UK at launch while Business Premium reportedly has them across all supported regions. Taken together with the cache-read halving described below, DevDay moved the price of tokens down and the price of capacity up.

TL;DR: GPT-6.1 Sol shipped at OpenAI DevDay on 29 September 2026 at $2 / $10 per million tokens. Claude Sonnet 5.5 shipped 28 September at $2 / $10 per million tokens. Identical headline, one day apart — and the meters underneath disagree on almost everything. Cached input is $0.10 on Sol against $0.20 on Sonnet 5.5. Sol charges 2x input and 1.5x output above 272K tokens; Anthropic charges the same rate across the full 1M window. On Artificial Analysis, Sol at max effort scores 52 for $0.72 per task while Sonnet 5.5 scores 56 for $7.60 — four index points apart, more than 10x apart on cost, driven by output volume rather than any published rate. Sol’s own cost per task moves 5.5x across effort levels with the rate card unchanged. And GPT-6 Sol was replaced seven days after launch, without being deprecated. The list price has become the least informative number on a model page.

Two vendors, one number, twenty-four hours

On 28 September Anthropic released Claude Sonnet 5.5 at $2 per million input tokens and $10 per million output. On 29 September, on stage at DevDay, OpenAI released GPT-6.1 Sol at $2 per million input tokens and $10 per million output.

The coincidence is not one. $2/$10 is now the contested mid-tier price point, the tier where the volume is, and both vendors have concluded independently that the way to win it is not to undercut. Anthropic held Sonnet 5.5’s rate card exactly flat against Sonnet 5 and put the improvement into speed and tool-call efficiency. OpenAI set GPT-6 Sol at $2/$10 a week ago and kept GPT-6.1 Sol at the same figure while cutting the cache rate underneath it.

When two competitors converge on an identical published price, that price stops carrying information. Everything that distinguishes the two products moves into the modifiers — and the modifiers are where these two models look nothing alike.

The meters underneath

GPT-6.1 SolClaude Sonnet 5.5
Released29 September 202628 September 2026
Input / output per MTok$2 / $10$2 / $10
Cached input / cache read$0.10 (0.05x)$0.20 (0.1x)
Context window1,050,0001,000,000
Long-context surcharge2x input, 1.5x output above 272KNone across the full window
Max output (sync)128K128K
Batch discountNot advertised50% on input and output
Knowledge cutoff30 April 2026June 2026
Effort levelslow, medium (default), high, xhigh, maxlow → max, default high

Three of those rows can each swing a real bill further than the headline price ever could.

Cache reads differ by 2x. GPT-6.1 Sol reads cached input at $0.10 per million — 5% of base input, and exactly half what GPT-6 Sol charged seven days earlier. Sonnet 5.5 sits on Anthropic’s standard 0.1x multiplier at $0.20. For the workload shape that dominates production agent traffic — a large stable system prompt, a fixed tool schema, retrieved documents replayed turn after turn — cache reads are the biggest line on the invoice. Halving them is a bigger intervention than anything either vendor did to the headline rate, and neither put it in a headline.

They disagree about whether long context costs extra. OpenAI applies 2x input and 1.5x output rates above 272K input tokens, the same cliff this desk mapped when GPT-6 Astra launched. Anthropic’s pricing page takes the opposite position in plain language: models from Claude 4.6 onward “include the full 1M token context window at standard pricing,” and “a 900k-token request is billed at the same per-token rate as a 9k-token request.” So a workload that routinely pushes past 272K — whole-repository analysis, long document review, deep agent transcripts — is priced identically at the top of the page and very differently at the bottom of it. Both models advertise roughly a million tokens of context. Only one of them sells all of it at the advertised price.

Anthropic has a batch lane. A 50% discount on both input and output for asynchronous work, bringing Sonnet 5.5 to $1/$5. No equivalent is advertised for GPT-6.1 Sol. For anything that does not need to answer in real time, that alone inverts the cache-read advantage.

The number that should end rate-card budgeting

Artificial Analysis has now scored both models, and the result is the clearest available argument that per-token price no longer predicts spend.

Model / effortIntelligence IndexCost per task
GPT-6.1 Sol (low)42$0.13
GPT-6.1 Sol (medium)48$0.21
GPT-6.1 Sol (high)50$0.32
GPT-6.1 Sol (xhigh)51$0.39
GPT-6.1 Sol (max)52$0.72
Claude Sonnet 5.556$7.60

Read the Sol column first, because it makes the point without any cross-vendor caveats attached. One model, one rate card, one unchanged pair of published prices — and cost per task moves 5.5x, from $0.13 to $0.72, purely on the effort setting. Nine index points of capability are available for roughly five and a half times the money, and nothing on the pricing page hints at the range.

Then the cross-vendor row. Sonnet 5.5 scores four points higher than Sol at max effort and costs more than ten times as much per task. That gap is not a fee schedule; it is tokens emitted. Artificial Analysis flags Sonnet 5.5 as “very verbose,” logging roughly 410M output tokens across its suite against an 81M median for comparable models, and output bills at $10 per million on both platforms.

Fairness requires three qualifications, and they matter. This is one harness running particular configurations, not a universal result. Sonnet 5.5’s API default effort is high, and Anthropic’s own documentation tells agentic users to start at medium — so the measured figure reflects a setting the vendor itself steers away from for tool-heavy work. And the Intelligence Index was repriced and restructured this month, so cross-version comparisons need care.

None of that rescues the rate card. Whichever way the caveats fall, two models at an identical published price produced per-task costs an order of magnitude apart on the same evaluation. That is a decisive result about the pricing page, whatever it turns out to be about the models.

Seven days from launch to superseded

The other thing worth recording from DevDay has nothing to do with price.

GPT-6 Sol launched on 23 September. It was patched on 26 September for a vision defect. It was superseded on 29 September. Seven days, launch to replacement — a point Artificial Analysis made on the night.

It has not been deprecated. Its documentation page carries no shutdown date and no retirement commitment; it simply says to see GPT-6.1 Sol for the newer Sol model. So every team with gpt-6-sol in a configuration file is untouched, still billing at $2/$10, still working — and now running a model the vendor has stopped treating as current, with a patch history it cannot reconstruct, because neither GPT-6 model exposes a dated snapshot to pin against.

That was the concern in this desk’s 26 September coverage, and it took three days to become concrete. Compare the other side of the same week: Anthropic marked Sonnet 5 “legacy” in the same breath as launching Sonnet 5.5, published a retirement commitment of no sooner than 30 June 2027, and wrote a migration guide enumerating what breaks. Both vendors superseded a model. Only one of them gave the older model a status and a clock — and OpenAI’s own legacy completions endpoint is separately running out of models with contradictory dates in its docs, which is the same lifecycle-hygiene gap in a different place.

The shape of this market now

Three same-price releases inside eight days is a pattern, not a run of coincidences. Fireworks matched Kimi K3’s price and put the discount in the token count. Anthropic’s Opus 5.5 delivered a 40% cost reduction that turned out to be a change in default effort. Now two flagship mid-tier models land on the identical figure twenty-four hours apart.

The published price has become a positioning signal — a claim about which tier a model belongs to — rather than a prediction of spend. The actual bill is set by four things that live in footnotes: the cache multiplier, the long-context surcharge, the batch lane, and the effort setting. Only the first three are contractual. The fourth is a runtime decision your own code makes, on most platforms by default, and on the evidence above it moves cost further than the other three combined.

For buyers that is a change in method, not just in arithmetic. A model selection made by comparing rate cards was defensible eighteen months ago. This week it is a coin flip dressed as analysis.

What to do

Price a real task, not a million tokens. Take a representative slice of your own traffic — a hundred production requests with their actual prompt sizes, cache-hit rates and tool loops — and run it end to end on both models at two or three effort levels. Compare total spend against output quality. It costs an afternoon and produces the only number that transfers to your invoice.

Find your cache-hit ratio before you choose. If most of your input tokens are cache reads, GPT-6.1 Sol’s $0.10 against Sonnet 5.5’s $0.20 is close to halving your dominant cost line. If your cache hit rate is low, that advantage largely evaporates and the other rows decide it.

Check your p99 prompt length against 272K. Not your average — the tail. A workload whose typical request is 40K tokens but whose monthly repository sweep hits 600K will look cheap in a pilot and bill at 2x input on the runs that matter most. Anthropic has no such threshold; OpenAI’s is a hard cliff.

Set effort explicitly, everywhere, on both platforms. Sol defaults to medium, Sonnet 5.5 defaults to high, and the setting moves cost per task several-fold. An unset effort parameter is a cost decision made by a default you did not choose. Developers shipping agent loops should treat it with the same care as a model ID.

Pin what you can, and diary what you cannot. Anthropic gives Sonnet 5 a legacy label and a June 2027 retirement date; put it in the calendar. OpenAI gives GPT-6 Sol neither, so the calendar entry has to be yours — a standing review of whether the model string in your configuration is still the one the vendor considers current.

Re-check any comparison you made on headline rates. Anyone weighing ChatGPT against Claude, shortlisting coding tools, or choosing between ChatGPT and Claude for an agent platform this quarter compared two numbers that are now identical. The decision has to be made on the footnotes, because the headline has stopped distinguishing anything.

Frequently asked questions

Are the two models really the same price?

On the headline line, to the cent. OpenAI's model page lists GPT-6.1 Sol at $2 per million input tokens and $10 per million output tokens. Anthropic's pricing page lists Claude Sonnet 5.5 at $2 per million base input tokens and $10 per million output tokens. Both carry a roughly 1M-token context window — 1,050,000 for Sol, 1,000,000 for Sonnet 5.5 — and both cap synchronous output at 128K tokens. Below that line they stop matching. Cached input on GPT-6.1 Sol is $0.10 per million, which is 5% of base input and half of what GPT-6 Sol charged a week earlier; cache hits on Sonnet 5.5 are $0.20 per million, the standard 0.1x multiplier. Sol applies 2x input and 1.5x output rates above 272K input tokens; Anthropic states the opposite explicitly, that models from Claude 4.6 onward 'include the full 1M token context window at standard pricing' and that 'a 900k-token request is billed at the same per-token rate as a 9k-token request.' Anthropic also offers a 50% Batch API discount. So the same two numbers sit on top of materially different meters.

Where does the 10x difference in cost per task come from?

From output volume, which is the meter that actually moves. Artificial Analysis scores GPT-6.1 Sol at max effort at 52 on its Intelligence Index for $0.72 per task, and Claude Sonnet 5.5 at 56 for $7.60 per task. Four index points apart, more than ten times apart on cost. The mechanism is not a hidden fee — it is tokens emitted. Artificial Analysis flags Sonnet 5.5 as 'very verbose,' recording roughly 410M output tokens across its evaluation suite against an 81M median for comparable models, and output bills at five times input on both platforms. Two caveats matter before anyone quotes that ratio in a budget. The comparison is one harness at particular configurations, and Sol's own cost per task swings from $0.13 at low effort to $0.72 at max — a 5.5x range on an unchanged rate card. And Sonnet 5.5's figure reflects a model whose API default effort is `high`. The honest reading is not that one model is ten times cheaper; it is that effort and verbosity dominate the bill so completely that the per-token price cannot predict it.

GPT-6 Sol is a week old. Why is there already a GPT-6.1?

GPT-6 Sol arrived on 23 September 2026 at $2/$10, was patched on 26 September for a vision defect, and was superseded by GPT-6.1 Sol on 29 September — seven days from launch to replacement, as Artificial Analysis noted on the night. GPT-6 Sol is not retired: its documentation page carries no deprecation notice and no shutdown date, and simply points readers to 'GPT-6.1 Sol for the newer Sol model.' That matters because of a gap this desk flagged three days before the replacement landed — neither current GPT-6 model exposes a dated snapshot ID to pin against. A team that wrote `gpt-6-sol` into its configuration has not been moved, but it is now running a model the vendor has quietly stopped treating as current, with a patch history it cannot reconstruct and no version string that distinguishes pre-fix from post-fix behaviour. The upgrade to `gpt-6.1-sol` is a deliberate act, which is the good part; the problem is that staying put is now also a decision, and nothing in the platform surfaces it as one.

Which one should a team actually pick?

The list price is the wrong input, so start with workload shape. If your prompts are long and heavily reused — a large fixed system prompt, retrieved documents, a stable tool schema — GPT-6.1 Sol's $0.10 cache reads are half Sonnet 5.5's $0.20, and on a cache-dominated bill that is close to a straight 50% cut on the largest line item. If your prompts routinely exceed 272K input tokens, the direction reverses hard: Sol charges 2x input and 1.5x output above that threshold while Sonnet 5.5 charges the same rate across the full window, so a single 600K-token request costs materially more on OpenAI. If your work is asynchronous, batch is not a differentiator: both vendors take 50% off through their batch APIs. If you need the most recent world knowledge, Sonnet 5.5's June 2026 cutoff beats Sol's 30 April 2026. And if raw capability per dollar on general reasoning is the criterion, the independent numbers favour Sol at every effort level tested. Almost none of that is visible from $2/$10.

What is the practical lesson for cost modelling?

Stop budgeting from the rate card and start budgeting from a measured task. Four multipliers now sit between the published price and the invoice, and all four are vendor-specific: the cache read multiplier (0.05x on Sol, 0.1x on Sonnet 5.5, and as low as 0.025x on Claude Fable 5.1), the long-context surcharge (2x/1.5x above 272K on Sol, none on Sonnet 5.5), the batch discount (50% on both), and the effort setting, which on Sol alone moves cost per task 5.5x without changing a single published number. The reliable method is to take a representative slice of your own traffic, run it end to end on both models at two or three effort levels, and compare total spend and output quality. That takes an afternoon and produces a number the pricing pages cannot. Anyone who approved a model on the strength of matching headline rates this week approved it on the least informative figure available.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.