Cognition's SWE-2 is 64% cheaper than a number it published for someone else — and it loses the one benchmark Devin is sold for
TL;DR: Cognition shipped SWE-2 into Devin on 10 September 2026 — its first model post-trained on a multi-trillion-parameter base, Moonshot’s Kimi K3 at 2.8T parameters — with three configurable effort levels produced inside a single RL run. The capability jump over SWE-1.7 is large and real: 73.0% on DeepSWE 1.1 (SWE-1.7: 37.7%), 92.8% on Terminal-Bench 2.1, 50.0% on FrontierCode 1.1 Main against Fable 5.1’s 50.9%. The claim is 64% cheaper at that score, “up to 70%” on X. The catch is that every dollar figure in the launch belongs to a competitor — Fable 5.1 Medium at $3.28 per FrontierCode task, Max at $12.83 — while SWE-2 has no rate card, no standalone API, no model card and no published context window, because it ships only inside Devin’s ACU meter. Back out the percentage and you get ~$1.18/task, which is an inference, not a price. And the row SWE-2 loses is the one that counts: Terminal-Bench 4, 27.3% against 55.8% and 57.9% — long-horizon agentic work, which is what Devin is sold for. For buyers: pilot on the free month, measure ACUs per completed task, and treat the benchmark table as vendor-run and unreplicated.
A real model, priced in a currency nobody published
Start with the part that is not in dispute. SWE-2 is a substantial piece of work and a bigger jump than SWE-1.7 was in July.
Cognition took Kimi K3, Moonshot’s 2.8-trillion-parameter open-weights base that had already been reinforcement-trained for agentic coding, and pushed its own RL stack on top — the first time it has done that at multi-trillion scale. The result adds five to six points over the base on many benchmarks, which suggests the K3 frontier was not exhausted by its original training. SWE-2 is also Cognition’s first model with configurable reasoning effort at medium, high and max, and the interesting engineering claim is that all three points on the cost-performance curve came out of one reinforcement learning run shaped by a cost penalty, rather than three separately trained variants.
Against its own predecessor the numbers are emphatic: SWE-2 medium matches or beats SWE-1.7 on FrontierCode while using 58% fewer turns and costing 81% less on average. Fewer turns is the honest efficiency metric for an agent, and it moved a long way.
Then the launch reaches the price, and the price is not there.
The only dollar figures are someone else’s
Cognition’s claim is that SWE-2 hits 50.0% on FrontierCode 1.1 Main — within a point of Claude Fable 5.1 at 50.9% — while costing 64% less to run at that score. To make that concrete, it publishes what the comparison costs: Fable 5.1 Medium at $3.28 per FrontierCode task, Fable 5.1 Max at $12.83.
It publishes nothing equivalent for SWE-2. No per-token rate. No per-task figure. No standalone API. No model card, and no stated context window.
You can back the number out — 64% below $3.28 implies roughly $1.18 per task — and that arithmetic is fine as far as it goes. It is an inference from a percentage, and it is not a price you can put in a budget, because SWE-2 is not sold as a model. It is a capability inside Devin, and Devin bills in ACUs — Agentic Computing Units, Cognition’s normalised bundle of VM time, model inference and networking, historically around fifteen minutes of autonomous work apiece.
That gap matters more than it first appears. A cheaper model reduces the inference share of an ACU. It does not automatically reduce how many ACUs your task consumes, and the ACU is what the invoice counts. Cognition has reported an internal unit-cost improvement — genuine, and good engineering — and has not said how much of it reaches a customer. That is the same shape of problem as Cursor’s always-on agent billing, where the spend controller sits on the vendor’s side of the meter.
The row it loses is the row Devin is sold for
| Benchmark | SWE-2 | Fable 5.1 | GPT-6 Astra |
|---|---|---|---|
| FrontierCode 1.1 Main | 50.0% | 50.9% | 53.3% |
| DeepSWE 1.1 | 73.0% | 67.4% | 74.1% |
| Terminal-Bench 2.1 | 92.8% | 91.4% | 89.9% |
| Terminal-Bench 4 | 27.3% | 55.8% | 57.9% |
Three of those four rows support the near-frontier story. The fourth does not, and it is not a random miss.
Terminal-Bench 4 is the harder successor benchmark, built around long-horizon, multi-step terminal work. SWE-2 scores less than half what either frontier model manages. The shape across the whole table is consistent: strong to excellent on well-scoped tasks, sharply weaker as the horizon lengthens.
That shape is awkward for this specific product. Devin’s pitch has always been asynchronous delegation — hand it a multi-hour task and walk away — which is exactly the regime Terminal-Bench 4 probes. And it interacts badly with ACU billing, because a failed long run still burns ACUs. A model that is 64% cheaper per attempt and needs two or three attempts on your hardest tickets is not cheaper. The cost claim and the weak row are measuring the same system from opposite ends, and only one of them made the headline.
To be fair to Cognition: SWE-1.7 scored 7.6% on the same benchmark. 27.3% is a 3.6x improvement. The trajectory is right; the level is not yet competitive.
Vendor-run, unreplicated, best-of-effort
No third party has reproduced any row of that table.
Two methodology notes belong next to it. Cognition states it reports the best score across reasoning-effort settings for every model. Applied evenly, that is fair — and it means the table compares best cases, while your team will run defaults. Second, each model was scored in its native harness (Claude Code for Anthropic’s models, the Devin CLI for others), and where a public result already existed, Cognition used it instead of running its own. The announcement does not mark which cells are internal and which are published, which makes independent verification harder than it needs to be.
This is not an accusation of bad faith; it is the ordinary condition of vendor benchmarks, and the reason private test sets and repriced index runs exist. The same caution applies here that applied when evaluation gaming showed up in METR’s work on GPT-5.6 Sol: a launch table is a hypothesis about your workload, not a measurement of it.
The base model has a week it did not ask for
One coincidence is worth naming plainly rather than leaving for someone else to discover.
SWE-2 is post-trained from Kimi K3. On the same day SWE-2 launched, Anthropic published a threat report naming Moonshot, K3’s maker, among seven China-based labs it accuses of illicit distillation — including the allegation that Moonshot relayed live Kimi customer requests to Claude without telling customers.
Two things are true simultaneously, and conflating them would be wrong in both directions.
Devin’s data path is not implicated. Cognition serves SWE-2 itself. K3’s weights are published. Running published weights relays nothing to Moonshot or anyone else — that distinction is the whole point of open weights.
The provenance chain is now contested, and a commercial buyer at the end of it has a fair indemnity question: if a base model’s capabilities are later held to be the product of a breach of a third party’s terms, who covers the downstream user? Anthropic’s allegations are vendor-authored, unreplicated and undenied. The proportionate response is a contract clause and a written answer from Cognition, not a veto — the same conclusion this desk reached about licensed model-factory arrangements when continuity, not capability, was the exposure.
What to do with this
- Pilot on the promotion, and measure ACUs per completed task. Not per attempt, not per benchmark point. Completed tasks are what you buy; everything else is what you are shown.
- Split your evaluation by task horizon. Well-scoped ticket work is where SWE-2’s near-Fable scores are most credible. Multi-hour autonomous delegation is where Terminal-Bench 4 predicts trouble, and where retries convert a cheap model into an expensive one.
- Run defaults, not best-of-effort. The launch table reports each model’s best effort setting. Your production config will not be that, for any vendor on the sheet.
- Ask what happens when the free month ends. A model with no published price, inside a meter you do not control, is a renegotiation that has not happened yet. Get the post-promotional cost in writing before you move a team onto it.
- Keep a frontier fallback wired in. If your hardest work is the long-horizon kind, the honest read of Cognition’s own numbers is to keep paying Fable 5.1 or GPT-6 Astra for it and route the rest to SWE-2.
- File the provenance question. One email to Cognition about base-model indemnity, answered in writing, costs nothing and closes a gap your procurement team will otherwise find later.
For the wider field, the Devin review tracks where this sits in the agent stack, best AI coding tools covers the alternatives, and the Claude Code and Cursor reviews cover the two harnesses SWE-2 is priced against. Cognition’s own trajectory — a $26B valuation and an ARR line built on Devin — is the context for why it keeps training its own models rather than renting someone else’s.
Frequently asked questions
What is SWE-2 and what did Cognition actually claim?
SWE-2 is Cognition's second in-house coding model, announced on 10 September 2026 and shipped into Devin Desktop and the Devin CLI immediately, with Devin Web and Devin Fusion rolling out behind it. It is post-trained from Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weights base that had already been reinforcement-trained for agentic coding, which makes it Cognition's first push of its own RL stack into the multi-trillion-parameter regime. It is also Cognition's first model with configurable reasoning effort — medium, high and max — and the notable engineering claim is that the whole cost-performance curve was optimised inside a single reinforcement learning run using a cost penalty, rather than trained as three separate models. The headline claim is that SWE-2 scores 50.0% on FrontierCode 1.1 Main, within one point of Claude Fable 5.1 at 50.9%, while costing 64% less to run at that score; Cognition's post on X broadened that to 'up to 70% lower cost'. Against its own predecessor the improvement is not marginal: SWE-2 medium matches or beats SWE-1.7 on FrontierCode while taking 58% fewer turns and costing 81% less on average.
Why is the 64% cost claim hard to use?
Because it is a ratio with only one side published. Cognition anchors the comparison to competitors' dollar figures — Fable 5.1 Medium at $3.28 per FrontierCode task and Fable 5.1 Max at $12.83 — and never publishes what SWE-2 costs per task or per token. Work the claim backwards and the implied figure is roughly $1.18 per task, but that is an inference from a percentage, not a rate card. There is no standalone API, no published per-token price, no model card and no stated context window, because SWE-2 is not sold as a model at all. It is a capability inside Devin, and Devin bills in ACUs — Agentic Computing Units, Cognition's normalised measure bundling VM time, inference and networking, historically around fifteen minutes of autonomous work each. A cheaper model reduces the inference component of an ACU. It does not necessarily reduce the number of ACUs a task consumes, and the customer meter runs on ACUs. So the honest statement is that Cognition has reported an internal cost improvement, which is a genuine engineering result, and has not told buyers how much of it reaches the invoice.
Which benchmark does SWE-2 lose, and does it matter?
Terminal-Bench 4, and it matters more than any row it wins. SWE-2 scores 27.3% there against Fable 5.1's 55.8% and GPT-6 Astra's 57.9% — less than half its frontier competitors on the benchmark specifically designed for long-horizon, multi-step terminal work. The rest of the sheet is strong: 50.0% on FrontierCode 1.1 Main against 50.9% and 53.3%, 73.0% on DeepSWE 1.1 against 67.4% and 74.1%, and 92.8% on Terminal-Bench 2.1, which beats both frontier models. But the shape of those results points somewhere specific. SWE-2 is at or near frontier on well-scoped tasks and falls off a cliff on tasks that require sustained autonomy — which is precisely the workload Devin is marketed for, and precisely the workload that makes a model cheap per task and expensive per outcome, because failed long runs still burn ACUs. A tool that is 64% cheaper per attempt and needs three attempts is not cheaper.
How independent are these numbers?
They are not. Every figure in the launch is Cognition's own, and no third party has replicated any row. Two methodology details deserve attention before the table is quoted anywhere that matters. Cognition states that for each model it reports the best score across reasoning-effort settings, which is applied evenly to competitors and is therefore fair, but which also means the table compares best cases rather than default configurations — and default configurations are what your team will actually run. Second, each model was scored in its native harness — Claude Code for Anthropic's models, the Devin CLI for others — and where a public result already existed Cognition used it rather than running its own. The announcement does not mark which cells are internal runs and which are published figures, which makes external verification difficult in a way that is probably incidental rather than deliberate but has the same effect. This is the same evaluation-hygiene problem that has now surfaced repeatedly this year: vendor-run benchmarks and buyer-run benchmarks diverge, and the divergence is always in the vendor's favour.
Does the Kimi K3 base model create a problem after Anthropic's threat report?
It creates a question worth asking, not a reason to avoid the product. SWE-2 is post-trained from Kimi K3, which is Moonshot AI's open-weights base model, and on 10 September 2026 — the same day SWE-2 launched — Anthropic published a threat report naming Moonshot among seven China-based labs it accuses of illicit distillation of Claude, including allegations that Moonshot relayed live Kimi customer requests to Claude without telling customers. Two things are true at once. First, none of that touches Devin's data path: Cognition serves SWE-2 itself, the weights are published, and running them relays nothing to Moonshot or anyone else. Second, the provenance chain behind the base model is now contested, and a commercial buyer standing at the end of that chain has a legitimate indemnity question — if a model's reasoning ability turns out to be the product of a breach of another vendor's terms, who covers the downstream user? Anthropic's allegations are vendor-authored and unproven, and Moonshot has not responded publicly. The sensible action is a clause, not a veto: ask Cognition what its indemnity covers regarding base-model provenance, and file the answer.
Should a team switch to SWE-2 this month?
Pilot it, do not standardise on it, and measure in ACUs rather than in benchmark points. The free promotional access reported across Devin's paid tiers makes the pilot close to costless, which is a good reason to run one and a poor reason to migrate a workflow. Three things to measure while the promotion lasts. Run your real task mix at each effort level and record ACU consumption per completed task, not per attempt — that is the only number that maps to your invoice and it is the number the launch does not contain. Watch failure behaviour on long, multi-step work specifically, because that is where the Terminal-Bench 4 result predicts trouble and where a cheap model gets expensive through retries. And note what happens when the promotion ends: a model with no published price inside a meter you do not control is a renegotiation you have not had yet. If your work is well-scoped ticket-shaped tasks, SWE-2 at near-Fable scores is genuinely attractive. If it is autonomous multi-hour delegation, the benchmark sheet says buy the frontier model and pay for it.
Sources
- Cognition — Introducing SWE-2: Pushing the Pareto Frontier (10 September 2026)
- CellCog — Cognition SWE-2: Benchmarks, the 64% Cost Claim, and the Row It Loses
- OfficeChai — Cognition releases SWE-2, says it performs close to frontier at 70% lower cost (11 September 2026)
- ExplainX — SWE-2: 50% FrontierCode, 64% cheaper than Fable 5.1
- TechTimes — Cognition SWE-2 beats frontier coding AI at 64% lower cost using single-run RL training (11 September 2026)
- Anthropic — Countering misuse of AI: September 2026 (threat report naming Moonshot, maker of the Kimi K3 base model)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.