AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Aug 20, 2026
·
grokspacexaimodelscoding

SpaceXAI ships Grok 4.6: it ties the GPT-5.6 Sol tier on the composite index, for a third of the price — but the coding benchmarks tell a narrower story

TL;DR: On 12 August 2026, SpaceXAI released Grok 4.6, a post-training upgrade over Grok 4.5 — same base model, more training. It scores 61 on the Artificial Analysis Intelligence Index (up from 56), tying GPT-5.6 Sol Max and ranking among the top handful of models, though still behind Fable 5. Pricing is unchanged at $2/$6 per million tokens (doubling above a 200K-token prompt), a fraction of the GPT-5.6 tier it now matches on the composite index. It ships with a new xhigh reasoning level, a 500K-token context, and — the real story — as the default model in Cursor. But the composite number hides the shape: Grok 4.6 leads on knowledge-work benchmarks while sitting last of four on Terminal-Bench and behind GPT-5.6 Sol Max on DeepSWE. Read it as a value play — near-frontier quality at a materially lower cost per task — not a new leader. A 42-second time-to-first-token is a real cost for interactive use. Benchmark it on your own work before migrating anything.

What shipped

SpaceXAI released Grok 4.6 on 12 August, roughly five weeks after Grok 4.5 went public. The framing matters: this is a post-training refresh, not a new base model. SpaceXAI held the foundation constant and spent the improvement on a longer supplemental training run, regenerated supervised fine-tuning across reasoning levels and domains, and reinforcement learning in agentic environments. One behavioural change stands out — the model now does more self-testing and verification, checking its own work on extended tasks.

The specs, confirmed across Artificial Analysis, MarkTechPost and SiliconANGLE:

The one number everyone will quote — and what it hides

Grok 4.6 scores 61 on Artificial Analysis’s composite Intelligence Index, up from 56 for Grok 4.5. That level ties GPT-5.6 Sol Max and, per VentureBeat, places it among the top three on the index — overtaking Moonshot’s Kimi K3 and closing most of the gap to the frontier. It still sits behind Fable 5, the ceiling model in Anthropic’s current lineup.

A single composite score is exactly the thing to distrust, though, because it averages away the shape of the gains. Break it apart and Grok 4.6 is two different models depending on the workload:

Artificial Analysis itself flags that some of the wins fall inside statistical confidence intervals — “statistical ties, not leads.” So the honest summary is: Grok 4.6 has closed the gap to the GPT-5.6 Sol tier on average, leads it on knowledge work, and still trails it where autonomous coding is measured.

The real story is distribution, not the benchmark

The benchmark line will get the headlines, but the more consequential fact is where Grok 4.6 runs. It is the default model in Cursor, one of the most widely used AI coding environments. That placement is not a neutral integration — SpaceX signed a $60 billion deal to acquire Cursor’s maker, Anysphere, in June 2026, and while that acquisition has not yet cleared regulatory review, Grok is already the model millions of Cursor users touch first.

That is how a value-tier model becomes a default without winning a benchmark war: not by beating Claude or GPT-5.6 on Terminal-Bench, but by being wired into the tool at a price the platform can afford to serve broadly. For an AI-tools buyer, the distribution question — what model does my IDE reach for by default, and can I change it — is now as important as the leaderboard.

Why this matters

The frontier is getting cheaper faster than it’s getting better. Grok 4.6 matches a tier that cost far more a few months ago, at $2/$6. The generational story across 2026 has been less about a new ceiling and more about the price of near-ceiling capability collapsing. For most buyers that is the more useful trend: the model you could not justify on cost last quarter is now a third of the price.

“Matches on the index” is not “matches on your task.” The split between Grok 4.6’s knowledge-work strength and its agentic-coding gap is the whole point. A composite score is a marketing artefact; your workload is not the composite. Teams doing research, analysis and drafting will find it genuinely frontier-adjacent; teams running long-horizon autonomous coding agents will still feel the distance to the top Claude and GPT-5.6 tiers.

Latency is the unglamorous catch. Grok 4.6’s time-to-first-token is about 42 seconds — far above the sub-three-second median for models in its price band — and output runs around 67 tokens per second, slightly below median. For batch and agentic jobs that is tolerable. For interactive back-and-forth it is a real friction, and it is the kind of number that never appears in a launch headline.

No open weights keeps this a rental. Unlike the open-weight tier led by Kimi K3, Grok 4.6 is API-and-product only. There is no self-hosting, no offline option, and no way to pin a version you control. For regulated or air-gapped environments that rules it out regardless of the benchmark.

Honest caveats

The benchmarks are early and partly vendor-adjacent. Artificial Analysis is an independent evaluator, which is why it anchors this piece, but several figures come from SpaceXAI’s own reporting on its API and some wins are statistical ties. Treat the composite score as directional, not decisive, and wait for a broader independent read on real tasks.

“Optimised for long-running agents” is a claim, not a measurement. The marketing positions Grok 4.6 for autonomous coding and multi-step work, yet the agentic-coding benchmarks are precisely where it trails. The self-verification behaviour may help in practice, but that is exactly the kind of improvement that needs independent, task-level testing to confirm.

Pricing has a cliff. The attractive $2/$6 rate doubles above a 200K-token prompt. Long-context work — the thing a 500K window invites — is billed at $4/$12, and per xAI’s own documentation the higher rate applies to all tokens in the request, not just those past the threshold. That materially changes the cost calculus for exactly the workloads the big context is meant to serve. It is also unresolved on Amazon Bedrock, whose model card publishes flat rates with no long-context band — see our breakdown of Grok 4.6’s Bedrock listing and its data-residency premium.

The verdict

Grok 4.6 is a clean, credible iteration: a five-point index gain, real progress on coding, unchanged pricing, and a smart new xhigh setting for hard problems. On the composite number it has pulled level with the GPT-5.6 Sol tier, and on knowledge work it is genuinely competitive with the frontier at a fraction of the cost.

But the composite number oversells it. This is a value and distribution play, not a new capability leader. It trails Fable 5 overall, trails GPT-5.6 Sol Max where autonomous coding is measured, sits last on Terminal-Bench, and carries a 42-second cold-start that interactive users will feel. Its biggest advantage is not a benchmark at all — it is being the default inside Cursor, at a price that makes broad deployment cheap.

Recommendation: if you already live in the SpaceXAI or Cursor orbit, or you run knowledge and analysis workloads where cost per task matters, Grok 4.6 is worth putting into your own evaluation now — it may be the best value at its capability level. If you run long-horizon autonomous coding, or you need open weights, offline use, or low interactive latency, the proven frontier options — the top Claude and GPT-5.6 tiers — remain the safer default. As always: benchmark on your own tasks before you move anything you depend on.

Frequently asked questions

How much better is Grok 4.6 than Grok 4.5?

On the composite Artificial Analysis Intelligence Index, Grok 4.6 scores 61 versus 56 for Grok 4.5 — a five-point lift that moves it from a near-frontier model to level with the GPT-5.6 Sol Max tier. The bigger jumps are on agentic coding: DeepSWE v1.1 rises 11.9 points to 65.9%, and Terminal-Bench v3.0 nearly doubles to 26%. It is a post-training upgrade on the same base model, so the gains come from a longer supplemental training run, regenerated fine-tuning data and reinforcement learning in agentic environments — not a larger model.

How much does Grok 4.6 cost?

For prompts under 200,000 tokens: $2 per million input tokens, $0.50 per million cached input, and $6 per million output — unchanged from Grok 4.5. Once a prompt crosses 200,000 tokens those rates double to $4, $1 and $12. A faster-serving variant is offered at roughly double the price. That headline $2/$6 is well below the GPT-5.6 Sol tier it now matches on the composite index, which is the model's main selling point.

Where can I use Grok 4.6?

Through the SpaceXAI API (model ID grok-4-6), in Cursor on all plans, as the default model in Grok Build, via routing on OpenRouter, Vercel and Cloudflare, and — since 19 August 2026 — on Amazon Bedrock in most AWS Regions. There is no open-weights release and no self-hosting option; it is API and product access only. The Bedrock listing is the significant addition for enterprises, because it removes the new-vendor objection, though the feature surface there varies by endpoint. The Cursor integration remains the most consequential for developers now that SpaceX's acquisition of Cursor's maker has closed.

Is Grok 4.6 better than Claude or GPT-5.6 for coding?

Not clearly, and it depends on the task. Grok 4.6 leads on knowledge-work benchmarks like GDPval-AA and AA-Briefcase and scores 69.9% on CursorBench, but it sits behind GPT-5.6 Sol Max on DeepSWE (65.9% vs 73%) and last of four listed models on Terminal-Bench. It trails Fable 5 on the overall index. For autonomous, long-horizon coding, the proven frontier options remain the top Claude and GPT-5.6 tiers; Grok 4.6 is best read as a strong cost-per-task option, especially inside Cursor.

Should I switch my workflow to Grok 4.6?

Only after benchmarking it on your own tasks. The value case is real if you are already in the SpaceXAI or Cursor orbit and care about cost per task. But the 42-second time-to-first-token is a genuine cost for interactive use, and the agentic-coding gaps mean it is not a drop-in replacement for the top Claude or GPT tiers on autonomous work. Test before you migrate anything you depend on.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.