Cognition launches SWE-2, a Kimi K3 post-trained coding model rivaling Fable 5.1 at 64% lower cost, with GPT-6 Astra powering Devin testing
Cognition released SWE-2, a proprietary 2.8T-parameter MoE coding model RL post-trained from Moonshot AI's Kimi K3, reporting 50.0% on FrontierCode 1.1 Main and 92.8% on Terminal-Bench 2.1 with Devin-only availability; a companion OpenAI case study details…
Reports from September 10–12, 2026 announce Cognition's SWE-2, its most advanced coding model, post-trained with reinforcement learning from Moonshot AI's 2.8T-parameter Kimi K3 base. SWE-2 is a proprietary mixture-of-experts model with 104B active parameters per token and vendor-reported scores of 50.0% on FrontierCode 1.1 Main, 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1, and 27.3% on Terminal-Bench 4.0, all pending independent replication. Sources agree SWE-2 is roughly one point behind Claude Fable 5.1 on FrontierCode 1.1 Main at a claimed 64% lower cost and beats Grok 4.6 and SWE-1.7, but they frame the comparison differently: one report says it matches Fable 5.1 and GPT-5.6 Sol, while another notes it trails Fable 5.1 and GPT-6 Astra by a wide margin on long-horizon Terminal-Bench 4.0 tasks. Cognition says it scaled RL to the multi-trillion-parameter regime for the first time, tripled RL environments, trained all three reasoning-effort levels in a single run using Pareto-informed, slope-matched cost penalties, and used NVFP4/FP8 quantization-aware training with speculative decoding; RL reportedly adds 5–6 points over the K3 base on many benchmarks, and SWE-2 medium takes 58% fewer turns and costs 81% less than SWE-1.7 on FrontierCode. SWE-2 is available now in Devin Desktop and CLI, with Web and Fusion rolling out; there are no open weights and no standalone per-token API, and it is free for paid tiers through October 10, 2026. In a companion OpenAI case study published September 11, 2026, Cognition describes integrating GPT-6 Astra across Devin's cloud agent, CLI, and desktop products to automate testing — including Devin testing the iPhone game Otter Run and returning a simulator recording plus a report of passed checks and untested areas, and fixing customer-reported bugs from screenshots with automatic verification screenshots — which co-founder Walden Yan says could reduce manual code review and raise shipping velocity.
- SWE-2 is a proprietary mixture-of-experts coding model with 2.8T total parameters and 104B active per token, RL post-trained by Cognition from Moonshot AI's 2.8T-parameter Kimi K3 base
- Vendor-reported benchmarks: FrontierCode 1.1 Main 50.0%, DeepSWE 1.1 73.0%, Terminal-Bench 2.1 92.8%, Terminal-Bench 4.0 27.3%; all figures pending independent replication
- SWE-2 is within 1 point of Claude Fable 5.1 on FrontierCode 1.1 Main at a claimed 64% lower cost, and beats Grok 4.6 and SWE-1.7
- Sources frame the competition differently: one report claims parity with Fable 5.1 and GPT-5.6 Sol, another says SWE-2 trails Fable 5.1 and GPT-6 Astra by a wide margin on long-horizon Terminal-Bench 4.0 tasks
- RL reportedly adds 5–6 points over the K3 base on many benchmarks
- First Cognition model with all three reasoning-effort levels trained in a single RL run using Pareto-informed, slope-matched linear cost penalties; RL environments tripled
- Serving stack uses NVFP4/FP8 quantization-aware training and speculative decoding
- SWE-2 medium takes 58% fewer turns and costs 81% less than SWE-1.7 on FrontierCode
Coverage timelineoldest first · each row is one article
- · 6d agoCognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Hacker News · AI· 72
Cognition released SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, near Fable 5.1 at 64% lower cost.
- · 6d agoCognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
Hacker News · AI· 62
Cognition releases SWE-2, a 2.8T-parameter MoE coding model post-trained from Kimi K3, scoring 92.8 on Terminal-Bench 2.1.
- · 5d agoCognition helps Devin test its own work with GPT‑6 Astra
OpenAI News· 36
Cognition integrates GPT-6 Astra into Devin, its CLI, and desktop products to automate testing and provide evidence for code review.
- · 4d agoCognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost
MarkTechPost· 58
Cognition released SWE-2, an RL post-trained coding model from Kimi K3, scoring 50.0% on FrontierCode 1.1 Main and available only inside Devin.