ZeroHour
Hacker News · AIpublished ()ingested cdnsteve
Part of a story covered by 8 sources: “GPT-6 Astra launch week: Critical cybersecurity threshold, benchmark leads, Cognition's SWE-2 challenge, and enterprise adoption by Devin and Perplexity” — merged summary and timeline →

Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

infoModel releaseimportance 62
AI summary · glm-5.3-flash

Cognition releases SWE-2, a 2.8T-parameter MoE coding model post-trained from Kimi K3, scoring 92.8 on Terminal-Bench 2.1.

SWE-2 is a proprietary mixture-of-experts model with 2.8T total parameters and 104B active per token, built on the Kimi K3 base with additional Cognition reinforcement-learning post-training for agentic coding. Vendor-reported benchmarks include FrontierCode 1.1 Main 50.0, DeepSWE 1.1 73.0, Terminal-Bench 2.1 92.8, and Terminal-Bench 4.0 27.3. It claims to be one point behind Claude Fable 5.1 on FrontierCode at a claimed 64% lower cost, but trails Fable 5.1 and GPT-6 Astra by a wide margin on long-horizon Terminal-Bench 4.0 tasks. The model is available today in Devin Desktop and CLI, with no published weights, no per-token API pricing, and all figures pending independent replication.

  • SWE-2 is a 2.8T-parameter MoE with 104B active per token, post-trained from Kimi K3
  • Self-reported benchmarks: FrontierCode 50.0, DeepSWE 73.0, Terminal-Bench 2.1 92.8, Terminal-Bench 4.0 27.3
  • Trails Claude Fable 5.1 and GPT-6 Astra on long-horizon Terminal-Bench 4.0 tasks
  • Proprietary weights, no per-token API; available in Devin Desktop and CLI
  • Serving stack uses NVFP4/FP8 kernels with quantization-aware training and speculative decoding
Full article459 words · extracted from tokenstead.ai · click to collapse

MoE premier

2.8T total params, 104B active per token (MoE) - the Kimi K3 base with Cognition’s post-training on top, and the first time Cognition has scaled RL into the multi-trillion-parameter regime. The base had already been RL-heavy for agentic coding; Cognition’s pass added another 5 to 6 points on most benchmarks.

  • Serving stack: MoE inference on NVFP4 and FP8 kernels with quantization-aware training; FP8 carries K, Q, V, and score computations in the MLA layers. A draft model retrained with SpecForge gives 15% longer accept lengths, and a prefill delayer lifts TPM per GPU and tokens/sec per request by 10 to 20% (TTFT takes the hit).
  • Effort levels: mean steps per run 53 (medium), 80 (high), 98 (max), against 127 for SWE-1.7. Medium posts a higher FrontierCode score than SWE-1.7 with 58% fewer turns and 81% lower average cost, and lands its first real edit at a median of step 18 (SWE-1.7: 48).

Bar chart: SWE-2 medium effort needs a mean of 53 steps per run versus 127 for SWE-1.7, and reaches its first edit at a median of step 18 versus 48

Benchmarks (Cognition self-reported): FrontierCode 1.1 Main 50.0, DeepSWE 1.1 73.0, Terminal-Bench 2.1 92.8, Terminal-Bench 4.0 27.3. The headline: 50.0 on FrontierCode is one point behind Claude Fable 5.1 (50.9) and 3.3 behind GPT-6 Astra (53.3) - at a claimed 64% lower cost than Fable 5.1 and a quarter of Astra’s. Terminal-Bench 2.1 is the highest number in the published table. The soft spot is Terminal-Bench 4.0, where SWE-2’s 27.3 trails Fable 5.1 (55.8) and GPT-6 Astra (57.9) by a wide margin - long-horizon agentic work is where the gap to the frontier still lives.

Scorecard of Cognition's launch table: SWE-2, SWE-1.7, Kimi K3, Grok 4.6, Fable 5.1, GPT-5.6 Sol and GPT-6 Astra on FrontierCode 1.1 Main, DeepSWE 1.1, Terminal-Bench 2.1 and Terminal-Bench 4.0. SWE-2 leads Terminal-Bench 2.1 at 92.8 but trails on Terminal-Bench 4.0 at 27.3

Bar chart of FrontierCode 1.1 Main scores: GPT-6 Astra 53.3, Fable 5.1 50.9, SWE-2 50.0, Grok 4.6 48.0, GPT-5.6 Sol 47.5, Kimi K3 44.2, SWE-1.7 42.0

Proprietary weights, no local run. Cognition has not published SWE-2 weights, so there is nothing to download and no quant ladder to wait for. It is available today in Devin Desktop and CLI, with rollout on Devin Web and Fusion. Cognition publishes no per-token API for SWE-2, so the cost-per-task comparisons (64% cheaper than Fable 5.1 at FrontierCode parity) are the pricing surface, not a $/1M rate card. Every figure here is Cognition’s own number, pending independent replication.

coding agentic

Parameters

2800.0B

License

proprietary

Developer

Cognition

Origin

🇺🇸 USA

Released

Sep 2026

What people are building with SWE-2

Real demos from X

Launch post: SWE-2, post-trained from Kimi K3 - FrontierCode 1.1 Main 50.0, Terminal-Bench 2.1 92.8 View on X →

Benchmark scores

Vendor-reported - from the developer's own model card / tech report

DeepSWE

73.0

FrontierCode 1.1 Main

50.0

Terminal-Bench 2.1

92.8

Terminal-Bench 4.0

27.3

Vendor-reported - from the developer's own model card / tech report

Or run it in the cloud

No per-token API provider pricing tracked for SWE-2 yet. For flagship list prices, see the calculator.

PRICE HISTORY

Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

Loading price history...

Text extracted automatically; images, tables and formatting may be missing. Original: https://tokenstead.ai/models/swe-2