ZeroHour
Story · 3 sources · 4 articlesfirst updated ()1

Cognition launches SWE-2, a Kimi K3 post-trained coding model rivaling Fable 5.1 at 64% lower cost, with GPT-6 Astra powering Devin testing

infoModel releaseimportance 72
What's new: Expanded the previously truncated companion-reporting line into full detail on the OpenAI case study (published 2026-09-11): Cognition integrates GPT-6 Astra across Devin's cloud agent, CLI, and desktop products to automate testing, including the Otter Run iPhone game example (simulator recording plus coverage report) and screenshot-based bug fixing with automatic verification, which Walden Yan…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

Cognition released SWE-2, a proprietary 2.8T-parameter MoE coding model RL post-trained from Moonshot AI's Kimi K3, reporting 50.0% on FrontierCode 1.1 Main and 92.8% on Terminal-Bench 2.1 with Devin-only availability; a companion OpenAI case study details…

Reports from September 10–12, 2026 announce Cognition's SWE-2, its most advanced coding model, post-trained with reinforcement learning from Moonshot AI's 2.8T-parameter Kimi K3 base. SWE-2 is a proprietary mixture-of-experts model with 104B active parameters per token and vendor-reported scores of 50.0% on FrontierCode 1.1 Main, 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1, and 27.3% on Terminal-Bench 4.0, all pending independent replication. Sources agree SWE-2 is roughly one point behind Claude Fable 5.1 on FrontierCode 1.1 Main at a claimed 64% lower cost and beats Grok 4.6 and SWE-1.7, but they frame the comparison differently: one report says it matches Fable 5.1 and GPT-5.6 Sol, while another notes it trails Fable 5.1 and GPT-6 Astra by a wide margin on long-horizon Terminal-Bench 4.0 tasks. Cognition says it scaled RL to the multi-trillion-parameter regime for the first time, tripled RL environments, trained all three reasoning-effort levels in a single run using Pareto-informed, slope-matched cost penalties, and used NVFP4/FP8 quantization-aware training with speculative decoding; RL reportedly adds 5–6 points over the K3 base on many benchmarks, and SWE-2 medium takes 58% fewer turns and costs 81% less than SWE-1.7 on FrontierCode. SWE-2 is available now in Devin Desktop and CLI, with Web and Fusion rolling out; there are no open weights and no standalone per-token API, and it is free for paid tiers through October 10, 2026. In a companion OpenAI case study published September 11, 2026, Cognition describes integrating GPT-6 Astra across Devin's cloud agent, CLI, and desktop products to automate testing — including Devin testing the iPhone game Otter Run and returning a simulator recording plus a report of passed checks and untested areas, and fixing customer-reported bugs from screenshots with automatic verification screenshots — which co-founder Walden Yan says could reduce manual code review and raise shipping velocity.

  • SWE-2 is a proprietary mixture-of-experts coding model with 2.8T total parameters and 104B active per token, RL post-trained by Cognition from Moonshot AI's 2.8T-parameter Kimi K3 base
  • Vendor-reported benchmarks: FrontierCode 1.1 Main 50.0%, DeepSWE 1.1 73.0%, Terminal-Bench 2.1 92.8%, Terminal-Bench 4.0 27.3%; all figures pending independent replication
  • SWE-2 is within 1 point of Claude Fable 5.1 on FrontierCode 1.1 Main at a claimed 64% lower cost, and beats Grok 4.6 and SWE-1.7
  • Sources frame the competition differently: one report claims parity with Fable 5.1 and GPT-5.6 Sol, another says SWE-2 trails Fable 5.1 and GPT-6 Astra by a wide margin on long-horizon Terminal-Bench 4.0 tasks
  • RL reportedly adds 5–6 points over the K3 base on many benchmarks
  • First Cognition model with all three reasoning-effort levels trained in a single RL run using Pareto-informed, slope-matched linear cost penalties; RL environments tripled
  • Serving stack uses NVFP4/FP8 quantization-aware training and speculative decoding
  • SWE-2 medium takes 58% fewer turns and costs 81% less than SWE-1.7 on FrontierCode

Coverage timeline

  1. · 6d ago
    Hacker News · AI· 72
    Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

    Cognition released SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, near Fable 5.1 at 64% lower cost.

  2. · 6d ago
    Hacker News · AI· 62
    Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

    Cognition releases SWE-2, a 2.8T-parameter MoE coding model post-trained from Kimi K3, scoring 92.8 on Terminal-Bench 2.1.

  3. · 5d ago
    OpenAI News· 36
    Cognition helps Devin test its own work with GPT‑6 Astra

    Cognition integrates GPT-6 Astra into Devin, its CLI, and desktop products to automate testing and provide evidence for code review.

  4. · 4d ago
    MarkTechPost· 58
    Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

    Cognition released SWE-2, an RL post-trained coding model from Kimi K3, scoring 50.0% on FrontierCode 1.1 Main and available only inside Devin.