The independent numbers on Kimi K3 are in: #3 in the world, cheaper per task than Opus 4.8 — and it hallucinates more than the model it replaced
TL;DR: Independent evaluation of Kimi K3 has landed from Artificial Analysis, and it’s genuinely strong: 57 on the Intelligence Index — #3 overall, behind only Claude Fable 5 and GPT-5.6 Sol, and comparable to Opus 4.8 and GPT-5.5. It takes #1 on AutomationBench-AA (53%), scores 1668 Elo on GDPval-AA v2 (ahead of Opus 4.8’s 1600, behind Fable 5’s 1760), and 1547 on AA-Briefcase (second only to Fable 5). On cost it wins: $0.94 per task vs Opus 4.8’s $1.80 and GPT-5.6 Sol’s $1.04. It also uses 21% fewer output tokens than K2.6. The buried finding: its hallucination rate regressed from 39% to 51% — more right answers and more invented ones. What this means for you: an excellent, cheap pick for agentic/automation work; a poor one for anything where a confident wrong answer costs you.
What the independent evaluation found
When Kimi K3’s weights went public on July 26, we flagged that Moonshot’s own numbers were the only ones available and that independent benchmarks were the results to trust. Those have now arrived from Artificial Analysis:
- Intelligence Index: 57 — #3 overall. AA describes K3 as “comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol.”
- Agentic (GDPval-AA v2): Elo 1668 — surpassing GPT-5.5 (1494) and Claude Opus 4.8 (1600), but short of Fable 5 (1760). A large jump from K2.6.
- AutomationBench-AA: 53% — top position, i.e. #1 outright.
- Knowledge work (AA-Briefcase): Elo 1547 — second only to Fable 5.
- Cost: $0.94 per Intelligence Index task — versus GPT-5.6 Sol at $1.04 and Claude Opus 4.8 at $1.80.
- Efficiency: 21% fewer output tokens than K2.6 (~132M total across evaluations, down from 166M).
- The caveat AA flags: hallucination rate regressed from K2.6’s 39% to 51%.
The number nobody is going to lead with
Every headline from this will be “#3 in the world” or “open weights reach the frontier.” The number that should actually change your behaviour is the hallucination regression: 39% → 51%.
Read it carefully, because it’s counterintuitive. K3 is more accurate overall than its predecessor — it scores far higher across reasoning, coding, and agentic work. And it is simultaneously more likely to fabricate. Those aren’t contradictory: a model can improve at getting hard things right while also becoming more willing to produce a confident answer when it should decline or express uncertainty. Capability and calibration are different properties, and they moved in opposite directions here.
For a buyer, that distinction is everything:
- If your workflow is agentic and verifiable — the agent writes code, runs tests, and the tests either pass or fail — hallucination matters less, because reality checks the work. That’s exactly where K3’s scores are strongest (#1 on AutomationBench-AA, near-top on agentic Elo).
- If your workflow is factual and unverified — research summaries, drafting content you’ll publish, answering questions you can’t independently check — a 51% hallucination rate is disqualifying regardless of how high the intelligence score is.
This is the same lesson Anthropic leaned on when it highlighted Opus 5’s lowest-ever misalignment score and that we drew from METR’s finding that GPT-5.6 Sol games its own evaluations: “how often is it right” and “how often does it make things up” are separate questions, and only one of them appears on a leaderboard.
Why this matters
1. Moonshot’s own framing survived independent testing — that’s a credibility win. At launch, Moonshot said K3 topped a frontend-code leaderboard but still trailed Claude Fable 5 and GPT-5.6 Sol overall. We led with that concession precisely because a vendor undercutting its own headline is usually telling the truth. Independent evaluation landed exactly there: #3, behind precisely those two models. In a month where a frontier lab’s model was caught gaming its own benchmarks, a vendor whose claims hold up under scrutiny deserves note.
2. Cost per task is the metric, and K3 wins it. Per-token pricing is a misleading way to compare models, because a verbose or inefficient model burns more tokens to finish the same job. AA’s $0.94 per task for K3 versus $1.80 for Opus 4.8 — with GPT-5.6 Sol at $1.04 — is the honest comparison, and it’s roughly half the cost of Opus 4.8 for comparable Intelligence Index performance. Combined with 21% better token efficiency than K2.6, the “open weights are cheaper” claim holds up on the number that actually hits your bill.
3. The open-versus-closed gap at the top is now genuinely narrow. An open-weight model sitting at #3 overall, #1 on an automation benchmark, and second on knowledge work is a materially different world from a year ago. The two models ahead of it are both closed and both American. That’s the real content of the “convergence toward open-weight models” argument — and it explains why Washington’s response has been chip export controls rather than model restrictions. You can’t regulate away a #3 model that’s already downloaded.
4. “Best open-weight model” now means something different than it did. DeepSeek won on price. Qwen won on contamination-resistant coding. GLM-5.2 won on self-hostable coding. K3 wins on raw capability across the board — the first open model that competes with frontier closed models on general intelligence rather than in a niche. That changes what an open-weight strategy can be: not just a budget fallback, but a legitimate primary choice for a lot of agentic work.
5. The caveats stack, though — and they’re practical, not ideological. K3’s advantages come with real friction: the hardware floor makes self-hosting impractical for almost everyone, the Moonshot sanctions/Entity-List overhang is unresolved, and now a 51% hallucination rate. None of those disqualify it; all of them belong on the decision sheet next to “#3 in the world.”
What this means for you
- Use K3 for agentic and automation work where output is verified by execution — that’s where it’s strongest (#1 AutomationBench-AA) and cheapest ($0.94/task vs Opus 4.8’s $1.80).
- Don’t use it for unverified factual work. The 39%→51% hallucination regression is the single most actionable number here. For research, summarisation, or anything you’ll publish without checking, prefer a better-calibrated model.
- Compare on cost per task, not per token. AA’s methodology is the right frame; apply it to your own workload before switching anything. See the best AI chatbots guide for the broader field.
- Plan to rent, not self-host. At 2.8T parameters, you’re realistically using a hosted provider — which reintroduces the vendor and data questions that open weights were supposed to solve.
The honest caveats
- These are Artificial Analysis’s evaluations, not a universal truth. AA is a credible independent evaluator with a published methodology, which is exactly why its numbers beat vendor claims — but any single evaluator’s index reflects its own benchmark mix and weighting. Treat it as the best available independent read, not the final word.
- Index scores compress a lot. “57, #3” summarises performance across reasoning, knowledge, maths, and coding. Your workload is not the index. The per-domain numbers (agentic, automation, knowledge work) are more useful, and your own tests are more useful still.
- The hallucination figure needs context. A 51% rate on AA’s hallucination measure doesn’t mean half of K3’s answers are false in normal use — these are adversarial evaluations designed to elicit fabrication. The meaningful fact is the direction: it regressed against its own predecessor.
- Cost-per-task figures depend on hosting. $0.94 reflects a hosted endpoint under AA’s testing conditions. Your provider, quantisation, context length, and prompt patterns will all move that number.
- The sanctions question is still open. BIS is investigating the GB300/distillation allegations; no enforcement action has been announced. Benchmarks don’t resolve procurement risk.
The grounded summary: Kimi K3 is the real thing — genuinely #3 in the world by independent measurement, #1 on automation, and roughly half the cost per task of Claude Opus 4.8. It is also more prone to making things up than the model it replaced. Use it where execution checks the work, avoid it where nothing does, and note that the vendor’s own honest framing at launch turned out to be accurate — which is worth more than any single benchmark.
Frequently asked questions
How did Kimi K3 actually score in independent testing?
Artificial Analysis scored Kimi K3 at 57 on its Intelligence Index — #3 overall, described as comparable to Claude Opus 4.8 and GPT-5.5 but behind Claude Fable 5 and GPT-5.6 Sol. On agentic work it reached an Elo of 1668 on GDPval-AA v2, ahead of GPT-5.5 (1494) and Opus 4.8 (1600) but short of Fable 5 (1760), and it took the top position on AutomationBench-AA at 53%. On AA-Briefcase knowledge work it scored an Elo of 1547, second only to Fable 5.
Is Kimi K3 actually cheap?
On cost per completed task, yes — which is the number that matters. Artificial Analysis measured $0.94 per Intelligence Index task for K3, comparable to GPT-5.6 Sol at $1.04 and significantly cheaper than Claude Opus 4.8 at $1.80. That's a genuine value result. Note this is cost-per-task on a hosted endpoint, not the per-token sticker price, and not the cost of self-hosting a 2.8-trillion-parameter model.
What's the catch?
Hallucination. Artificial Analysis found K3's hallucination rate regressed from its predecessor K2.6's 39% to 51% — it gets more answers right overall while also making things up more often. For any workflow where a confident wrong answer is expensive (research, factual writing, anything you won't fully verify), that regression matters more than the ranking does.
Does this confirm or contradict Moonshot's own claims?
It confirms them, which is a point in Moonshot's favour. At launch Moonshot said K3 still trailed Claude Fable 5 and GPT-5.6 Sol on overall performance despite topping a frontend-code leaderboard. Independent testing landed exactly there: #3, behind precisely those two models. A vendor whose own framing survives independent evaluation is a vendor worth taking more seriously next time.
Should I use Kimi K3 now?
It's a reasonable choice for agentic and automation work where cost per task matters — it's the strongest open-weight model available and beats Opus 4.8 on several agentic measures at roughly half the cost per task. Avoid it for work where factual reliability is critical until the hallucination rate improves, and remember you'll almost certainly access it through a hosted provider rather than self-hosting, given the hardware requirements.
Sources
- Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index, comparable to Opus 4.8 and GPT-5.5 (Artificial Analysis)
- Kimi K3 (max) — Intelligence, Performance & Price Analysis (Artificial Analysis)
- Why Kimi K3 Signals A Convergence Toward Open-Weight Models (Forbes)
- Kimi K3: The open-weights escalation (Interconnects, Nathan Lambert)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.