AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Sep 22, 2026
·
spacexaigrokpricingcodingbenchmarksprocurement

Grok 4.7 keeps the $2/$6 rate card and changes the thing the rate card doesn't measure — a bigger base model, trained to work longer

TL;DR: SpaceXAI released Grok 4.7 on 21 September 2026 at $2 input / $6 output per million tokens — unchanged from Grok 4.6, and the launch page says it plainly: “Served at the same price and speed as Grok 4.6.” Three sentences later it says the other half: 4.7 “uses a new, larger base model,” was trained on a mix “weighted toward problems that take many hours to complete,” and “works longer on difficult tasks.” The benchmark table agrees in a column header — 4.7 is measured at xHigh effort, 4.6 at High. On a per-token API, a flat rate card plus more tokens per task is not a flat bill. Scores: CursorBench 4.0 46.3% (Fable 5.1: 51.8%), DeepSWE v1.1 71.0%* where the asterisk means high effort, off-column, Terminal-Bench 4.0 38.0% (Fable 5.1: 57.9%), EEBench 64.0% (best in table), Harvey Legal 19.6% (next best: 6.7%), HealthBench Professional 56.7% (behind both rivals). GPT-6 Astra appears nowhere in the price table and only in the one chart where it finishes last. Live today in Cursor, Grok Build, the API and third-party harnesses.

What shipped

Grok 4.7 went live on 21 September 2026 in Cursor, Grok Build, the Grok API, third-party coding harnesses, and model routers and cloud platforms. SpaceXAI describes it as “our most capable model for coding and knowledge work” and leads the page with the marketing line “Twice as fast, at half the price of comparable models.”

The pricing is $2 per million input tokens and $6 per million output tokens, with a fast variant at twice the output speed for twice the price. That is the identical structure Grok 4.6 has carried since its August launch, and SpaceXAI flags the continuity itself: “Served at the same price and speed as Grok 4.6.”

The benchmark table, reproduced from the launch page:

Grok 4.7 (xHigh)Grok 4.6 (High)GPT-5.6 Sol (Max)Fable 5.1 (Max)
Input $/M$2$2$4$10
Output $/M$6$6$20$50
CursorBench 4.046.3%40.4%41.7%51.8%
DeepSWE v1.171.0%*65.2%72.7%70.0%
EEBench64.0%53.0%39.4%56.4%
AA Briefcase v1.11,6571,5461,4871,678
Terminal-Bench 4.038.0%20.3%37.3%57.9%
Harvey Legal Agent19.6%15.8%2.5%6.7%
HealthBench Professional56.7%48.5%60.5%62.1%

* high effort

Separately, a GDPval Elo chart puts Fable 5.1 (max) at 1,735, Grok 4.7 (xhigh) at 1,695, Grok 4.6 (high) at 1,605 and GPT-6 Astra (max) at 1,542.

The rate card is not the bill

The most important sentence in the release is not a price. It is this: Grok 4.7 “works longer on difficult tasks, checks its own work more carefully.” Followed by: it “was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete.”

Both sentences describe a model that emits more tokens to finish the same job. On an API billed per token, that is a cost disclosure written in the vocabulary of a capability claim. It is not deceptive — it is the accurate description of what was built — but it means “same price as Grok 4.6” is a statement about the rate card and not about the invoice, and only the rate card got published.

This is the arithmetic that has quietly become the defining variable of 2026 model pricing. It is the same structure as GPT-6 Astra’s launch, where the headline per-token figure told buyers very little about cost per finished task, and the same structure as OpenAI’s effort toggle on Astra, where the setting that controls reasoning spend became the setting that controls the bill. A token price is a unit price. Reasoning effort is the quantity. Labs publish the first one in bold and the second one in a column header.

To SpaceXAI’s credit, the launch page contains exactly the right chart: CursorBench 4.0 score plotted against average cost per task, with Fable 5.1, Opus 5, GPT-5.6 Sol, Sonnet 5 and Grok 4.7 positioned on it, and toggles for cost, tokens and steps. That is the honest measurement. It is also the only figure on the page whose axes carry no published values — the x-axis runs $0 to $18 in gridlines, and where each model actually sits is a dot, not a number.

The asterisk inside the asterisk

Look again at the column headers. Grok 4.7 is measured at xHigh. Grok 4.6 is measured at High. GPT-5.6 Sol and Fable 5.1 are at Max.

That makes the 4.7-versus-4.6 comparison a comparison of two models at two different spend settings. Terminal-Bench moving from 20.3% to 38.0% is a genuinely large jump; how much of it is the new larger base model and how much is one extra notch of reasoning effort is not something the table separates, and it is the question a buyer sizing a migration actually needs answered.

Then there is the footnote. DeepSWE v1.1 reads 71.0%*, and the asterisk resolves to “high effort” — one level below the xHigh heading its own column carries. No other cell in that column is marked. So the single most quoted number in the release, the one that sits within two points of GPT-5.6 Sol’s 72.7%, was produced at a different setting from its neighbours. Anyone repeating “Grok 4.7 hits 71% on DeepSWE” should carry the effort level along with it, because the comparison changes depending on which way that footnote cuts.

None of this is concealment. Every label is printed on the page. It is a reminder that in 2026 a benchmark score without an effort setting attached is an incomplete number, in the same way a context window without a price-per-token-above-the-threshold is incomplete.

The comparator that moved charts

The price-and-benchmark table compares Grok 4.7 against Grok 4.6, GPT-5.6 Sol and Fable 5.1. GPT-6 Astra — OpenAI’s current flagship, launched 3 September at $10 input and $50 output — is not in it.

Astra appears once, in the GDPval Elo chart, where it finishes last at 1,542, behind not only Grok 4.7 but Grok 4.6.

Both choices are defensible in isolation and instructive together. Comparing against GPT-5.6 Sol at $4/$20 rather than Astra at $10/$50 makes “half the price of comparable models” a much harder claim to make, so including Sol in the price table is arguably the conservative framing. But including Astra in precisely the one chart where it loses, and excluding it from the seven-row table where it might not, is comparator selection doing work. Read the table as what a vendor chose to show, then go get the rows it did not.

Where it is genuinely strong, and where it is not

Strong, and interesting. EEBench at 64.0% is the best score in the table by nearly eight points over Fable 5.1, on electrical engineering — a domain where none of these models has been a default. AA Briefcase v1.1 at 1,657 is within 21 points of Fable 5.1’s 1,678 at a fifth of the rate card. And the Harvey Legal Agent Benchmark result, 19.6% against GPT-5.6 Sol’s 2.5% and Fable 5.1’s 6.7%, is a three-to-one lead over the field.

That legal number deserves both attention and suspicion. It arrives days after OpenAI shipped a legal vertical built on the same partner layer Harvey occupies, which makes any Harvey-built benchmark commercially loaded in both directions. The more durable reading is the absolute level: every model in the table fails more than 80% of the tasks. A three-times lead on a benchmark where the leader scores 19.6% tells a law firm which model to pilot, not that the work is automatable.

Not strong. Terminal-Bench 4.0 at 38.0% against Fable 5.1’s 57.9% is a twenty-point gap on exactly the workload — long-running autonomous terminal work — where the cheap token rate is supposed to pay off. It does not, because an agent run that fails at hour three has consumed its full token cost and produced nothing. Cost per completed task and cost per token diverge hardest here, and this is the row where a buyer should be most sceptical of the headline price. HealthBench Professional at 56.7% sits behind both GPT-5.6 Sol (60.5%) and Fable 5.1 (62.1%), which is worth knowing before anyone routes clinical-adjacent work by price.

The safety stack, and a door that is not open to you

SpaceXAI says 4.7 ships “an entirely new safeguard stack,” calls it “the strongest model we’ve tested on refusals and jailbreak resistance,” reports 62.4% on LatchBio’s biosafety benchmark, and says HackerBench v0.3 shows it “allowing only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work.”

That last pairing is the genuinely useful metric design — a refusal rate is meaningless without the false-refusal rate beside it, and security teams have spent a year complaining that safety tuning blocks legitimate defensive work. If the 3.3% figure holds under outside scrutiny it is a real result.

The line after it is the one to file: “We’ve also started giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research.” This is now a standard shape. Capability that exists in the model is gated behind a named-partner programme, and the general rate card buys the restricted version. It is the same structure as the frontier capability tiers that spread across the industry this quarter: the price list is no longer the whole product catalogue, and access is a clearance, not a purchase.

What to do this week

Measure cost per completed task, not cost per token. Take one representative workload already running on Grok 4.6 or a rival, run it on 4.7, and compare total spend to finish. SpaceXAI charted this metric and did not label it; the number from an actual workload is the one that settles the question.

Pin the effort setting explicitly. Do not inherit a default. If the migration is evaluated at xHigh because that is what the launch table used, budget for xHigh. If the budget assumes High, evaluate at High — the published gains are partly a spend decision, and inheriting it silently is how a flat rate card produces a surprising invoice.

Weight the benchmark that matches the work. Supervised, editor-based coding: CursorBench, where 4.7 beats GPT-5.6 Sol and trails Fable 5.1 at a fifth of Fable’s price. Unattended multi-hour agents: Terminal-Bench, where the twenty-point deficit to Fable 5.1 is the whole story and the price advantage does not cover it. These two rows recommend opposite decisions for the same model, which is why a single “best model” verdict is the wrong output.

Do not re-rank on the launch table alone. Four effort levels across four columns, one footnoted cell, and a missing flagship comparator mean the table is a starting hypothesis. Independent evaluation on matched effort settings is what will make these numbers comparable, and it has not been published yet.

Check availability before assuming parity. Grok 4.7 is listed as live in Cursor, Grok Build, the API, third-party harnesses and cloud routers. Prior Grok releases reached regions and clouds on a lag, and the Bedrock rollout for 4.6 split data residency and endpoint pricing in ways the launch post did not mention. Verify the specific surface a team depends on.

The pattern

Update, 22 September 2026 — the index score has company, at a seventh of the output price. Artificial Analysis places Grok 4.7 at xHigh effort at 46 on Intelligence Index v4.3.2. One day after this launch, Xiaomi’s MiMo-V2.6-Pro arrived on the same score of 46 under an MIT licence, priced at $0.435 input / $0.87 output per million tokens against Grok 4.7’s $2/$6, with downloadable weights and the RL training code published alongside them. Index ties are fragile — different versions, different effort settings, different benchmark revisions — and this one should not be read as capability parity. What it does do is give the “twice as fast, at half the price of comparable models” framing a comparator the launch page did not anticipate, and it makes the cost-per-completed-task exercise below the only way to settle the question.

Update, 22 September 2026 — the fourth launch inverts the move. One day later Anthropic released Claude Opus 5.5 and ran this article’s mechanism backwards: rather than holding the rate card and raising the effort column, it cut the rate card 20% ($5/$25 to $4/$20, with cache reads down 60% to $0.20) and lowered the default effort from high to medium, then reported a headline saving of 40% “at default settings.” Same structural problem, opposite direction — the advertised number describes a configuration rather than a model, and the two halves of it are not separable from the price list. Opus 5.5’s reported 66.4% on Terminal-Bench 4.0 is also the first direct challenge to the reading below that Grok 4.7’s twenty-point Terminal-Bench deficit to Fable 5.1 is decisive: a model at $4/$20 now claims to beat the $10/$50 flagship on that benchmark, which changes the price-per-capability frontier that Grok 4.7’s $2/$6 was defending.

Three model launches this month have now used the same move: hold the per-token price, raise the reasoning spend, and report the benchmark at the higher effort setting. It is not a trick — the model really is better, and the price of a token really did not change. But it makes the most visible number on a launch page the least informative one, and it pushes the number that matters into a chart with unlabelled axes.

The buyer-side response is unglamorous and increasingly mandatory: stop comparing rate cards, start comparing invoices for a fixed unit of work, and treat any benchmark score without an effort level attached the way a price without a currency would be treated. Grok 4.7 looks like a genuinely good model at a genuinely aggressive price. Whether it is cheaper is a different question from whether it is cheap, and only one of those two the launch page answers.

For the broader picture of where Grok sits against the field, see the Grok review, the Grok vs ChatGPT comparison and the best AI coding tools roundup.

Frequently asked questions

Did the price of Grok actually go up?

The rate card did not move. Grok 4.7 is listed at $2 per million input tokens and $6 per million output tokens, exactly what Grok 4.6 has carried, with a fast variant at twice the output speed for twice the price. What is likely to move is tokens consumed per task, and SpaceXAI says as much in its own description: the model 'works longer on difficult tasks, checks its own work more carefully,' and was trained on a mix 'weighted toward problems that take many hours to complete.' Those are token-generating behaviours. A per-token API bills the behaviour, not the rate card. The honest way to read the launch is that the price of a token is unchanged and the number of tokens in a unit of work is not, and only one of those two numbers was published.

What does the xHigh label in the benchmark table change?

It changes what the comparison is comparing. In the launch table, Grok 4.7's column is headed xHigh, Grok 4.6's is headed High, GPT-5.6 Sol's is Max and Fable 5.1's is Max. Effort settings control how much reasoning a model spends before answering, so a score at xHigh and a score at High are not two measurements of the same thing — they are two measurements at two different costs. The 4.7-versus-4.6 gains (CursorBench 46.3% against 40.4%, Terminal-Bench 38.0% against 20.3%, EEBench 64.0% against 53.0%) are therefore a combination of a better model and a higher spend setting, in a proportion the page does not separate. Nothing here is hidden — the labels are printed — but a buyer reading the table as a like-for-like upgrade chart is reading it wrong.

Why does the DeepSWE score have its own asterisk?

Because it was measured at a different effort level from the rest of its own column. The table header puts Grok 4.7 at xHigh, but the DeepSWE v1.1 entry reads '71.0%*' with the footnote '* high effort' — one notch down. Every other cell in that column carries no asterisk. Read at face value, it means the DeepSWE figure is the only score in the column produced at a cheaper setting, which is unusual, since the normal direction of a footnote like this is to mark the one number that needed extra spend rather than less. Either way it is the single most-cited benchmark on the page and it is the one measured off-column, so anyone quoting 71.0% against GPT-5.6 Sol's 72.7% should quote the effort level with it.

Is Grok 4.7 now the right default for a coding team?

It is a credible candidate at the price and a poor fit for one specific workload. On CursorBench 4.0 it posts 46.3% against Fable 5.1's 51.8% and GPT-5.6 Sol's 41.7%, which at a fifth of Fable's rate card is a real value argument for interactive, shorter-horizon coding. On Terminal-Bench 4.0 — multi-hour autonomous terminal work — it posts 38.0% against Fable 5.1's 57.9%, a gap wide enough that the cheaper token rate does not close it, because a failed long-running agent run costs the whole run. The practical split is to treat 4.7 as a strong price-performance option for supervised coding inside an editor and to keep a more expensive model on unattended multi-hour agent jobs until the cost-per-completed-task numbers come in from independent evaluation rather than from the launch page.

What is going on with the legal benchmark result?

Grok 4.7 posts 19.6% on the Harvey Legal Agent Benchmark against GPT-5.6 Sol's 2.5% and Fable 5.1's 6.7% — roughly three times the next-best score in the table, on a benchmark built by a legal-AI vendor. That is the most striking single number in the release and the one most in need of outside confirmation, because a large lead on a narrow vendor-built evaluation is exactly the shape that either signals a genuine capability or signals fit to a particular task format. It also lands days after OpenAI shipped a legal vertical of its own, which makes the comparison commercially loaded. Absolute numbers under 20% on all four models are the more useful takeaway: on this evaluation every frontier model fails the large majority of tasks, and none of them is close to replacing the workflow.

What should a team do about this in the next week?

Three things, in order. Re-run an existing representative workload on 4.7 and compare total cost per completed task rather than cost per million tokens — that is the measurement the launch page charts and does not label. Pin the effort setting explicitly in the API call instead of inheriting a default, so the reasoning spend is a decision rather than a surprise. And check the two benchmarks that most resemble the work actually being done: teams doing supervised editor-based coding should weight CursorBench, teams running unattended agents should weight Terminal-Bench, and the two point in opposite directions for this model.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.