AInspiro
Tool Reviews

SWE-2 review: 0.9 points behind the frontier, 64 percent cheaper

AInspiro Editorial·
This article was created with AI assistance.

The 2026 coding-model story is not about who tops the chart

It is about who gets close enough, cheaply enough. Cognition released SWE-2 on September 10, the coding model inside its Devin agent. On its own FrontierCode 1.1 Main benchmark it scores 50.0 percent, just 0.9 points behind Anthropic's Fable 5.1 at 50.9 percent, and Cognition says it runs 64 percent cheaper than Fable 5.1.

The full leaderboard, not just the headline

On FrontierCode 1.1 Main: GPT-6 Astra 53.3, Fable 5.1 50.9, SWE-2 50.0, Grok 4.6 48.0, GPT-5.6 Sol 47.5, Kimi K3 44.2, and Cognition's own prior SWE-1.7 at 42.0. So SWE-2 did not overtake anyone. It planted itself right against the front cluster.

It also makes a bolder claim: on long-horizon agentic work it approaches GPT-6 Astra at roughly a quarter of the cost. For teams running coding agents all day, this is a budget question, not a vanity question.

The base model is the interesting detail

SWE-2 was not trained from scratch. It is post-trained from Moonshot's Kimi K3, a 2.8-trillion-parameter open model from China. A US startup takes a Chinese open base, post-trains it, and pushes it near the frontier. That alone is a footnote on how the open-weights supply chain now crosses borders.

On its medium effort setting it uses 58 percent fewer turns and costs 81 percent less than its own SWE-1.7, and reaches its first real code edit in a median 18 steps versus 48. Fewer wasted loops is the most immediately felt improvement in a coding agent. The capability also rolled into four Devin surfaces: Desktop, CLI, Web, and Fusion.

But do not read only that one row

In the same table, on Terminal-Bench 4, a much harder agentic benchmark, SWE-2 scores just 27.3 percent, while Fable 5.1 is 55.8 and GPT-6 Astra 57.9. That is not a rounding gap. It says SWE-2's price advantage is strongest inside the cleaner software-engineering lane, and once tasks get harder and longer, Anthropic and OpenAI are still far ahead.

And every number above is Cognition's self-reported result on its own benchmark, with FrontierCode built by Cognition and the cost comparisons calculated by Cognition. Independent evaluation and real-world use inside Devin will tell the fuller story. The model was not released as standalone weights or a public per-token price, but folded directly into Devin's product.

Verdict: the template for a good-enough strategy

SWE-2 does not win first place on any chart. What it wins is the judgment that ordinary engineering work is good enough and clearly cheaper. If your team is comparing Devin with Cursor, Copilot, or Claude Code, the question is not whether it takes some number one. It is whether it clears enough of your real work at a lower cost to make the trade worth taking.

Review advice: before renewing a frontier-model contract, run SWE-2 against Fable 5.1 and GPT-6 Astra on your own repositories. Keep long-horizon terminal tasks on another model until its Terminal-Bench 4 gap is tested on your tasks. Routing different difficulties to different tiers is steadier than betting on a single model.

A note on availability

SWE-2 is not sold as a standalone API or a downloadable weight. It lives inside Devin's product surfaces, which means the way most teams will meet it is through a Devin subscription rather than a direct model call. That bundling is part of the cost story: you are not buying tokens, you are buying a coding agent that happens to run this model. For buyers, the practical question becomes whether Devin's workflow fits your repo and review process, not whether the benchmark number is impressive in isolation.

What this means for you

You do not need the strongest model. The most cost-effective play in 2026 is routing work of different difficulty to different model tiers, and reserving the expensive ones for the long jobs that actually need them. A model that hugs the frontier while staying cheap is exactly the middle tier your routing strategy should be filling. Benchmark it in your own scenario before buying, and ignore the launch-day numbers.