ZeroHour
Story · 1 source · 1 articlefirst updated ()

Frontier model race: OpenAI's GPT-6 Astra launch now contested by Cognition's SWE-2 coding model

infoModel releaseimportance 94
What's new: Since the previous story summary (2026-09-08), the new development is Cognition's launch of SWE-2 (2026-09-10), a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, within one point of Claude Fable 5.1 at 64% lower cost, and matches Fable 5.1 and GPT-5.6 Sol at a fraction of the price — intensifying competition with GPT-6 Astra in the frontier coding space. The…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

Days after OpenAI launched GPT-6 Astra (claims: 99.9% ARC-AGI-3, ~98% FrontierMath, 100% ExploitBench; ~$6/hour agentic engineering), Cognition shipped SWE-2, post-trained from the 2.8T-parameter Kimi K3, scoring 50.0% on FrontierCode 1.1 Main — within one…

OpenAI launched GPT-6 Astra as its new flagship model, described by Latent Space as its first Stargate and lightly looped frontier model. Both Latent Space reports cite 99.9% on ARC-AGI-3, but the FrontierMath figures differ: Latent Space (2026-09-03) reports 97.6% on the hardest FrontierMath, while AINews (2026-09-04) reports OpenAI's claim of 98% on FrontierMath Tier 4; OpenAI also claims 100% on ExploitBench. Latent Space tested Astra with over 20 billion tokens, reporting it can train and select models, label data, deploy and debug systems, and orchestrate 20-50 parallel subagents, at roughly $6 per hour of agentic engineering (33 tokens per second), with token efficiency independently confirmed by Artificial Analysis. Pricing is $10/$50 per 1M input/output tokens standard ($20/$100 fast tier); rollout starts with limited organizations, then ChatGPT Plus/Pro/Business/Enterprise, the API, and AWS. Sources disagree on Astra versus Claude Fable 5.1: Latent Space reports Astra beats Fable 5.1 on many metrics, while Artificial Analysis's independent scores (67 Coding Agent Index, 61 Intelligence Index) place it behind; the system card notes improved alignment but decreased chain-of-thought monitorability. New since the previous story summary: on 2026-09-10 Cognition launched SWE-2, post-trained from the 2.8T-parameter Kimi K3 base model, scoring 50.0% on FrontierCode 1.1 Main (within one point of Fable 5.1 at 64% lower cost), 73.0% on DeepSWE 1.1, and 92.8% on Terminal-Bench 2.1, beating Grok 4.6 and SWE-1.7 while matching Fable 5.1 and GPT-5.6 Sol at a fraction of the price. Cognition says it scaled reinforcement learning to the multi-trillion-parameter regime for the first time, using Pareto-informed cost penalties that train all reasoning-effort levels in one run, tripled RL environments, and NVFP4/FP8 quantization-aware training; Kimi K3 post-training reportedly adds 5-6 points on many benchmarks, and SWE-2 medium takes 58% fewer turns at 81% lower cost than SWE-1.7 on FrontierCode. SWE-2 is available in Devin Desktop and CLI, with rollout on Devin Web and Fusion.

  • GPT-6 Astra: OpenAI claims 99.9% on ARC-AGI-3, 98% on FrontierMath Tier 4, and 100% on ExploitBench; Latent Space's earlier report cites 97.6% on the hardest FrontierMath (sources disagree on the exact FrontierMath figure).
  • Latent Space tested GPT-6 Astra with over 20 billion tokens: it can train and select models, label data, deploy and debug systems, and orchestrate 20-50 parallel subagents.
  • Latent Space measured about $6 per hour of agentic engineering at 33 tokens per second; Artificial Analysis independently confirmed the token efficiency.
  • GPT-6 Astra pricing: $10/$50 per 1M input/output tokens standard, $20/$100 fast tier; rollout order is limited organizations, then ChatGPT Plus/Pro/Business/Enterprise, the API, and AWS.
  • Sources disagree on Astra vs Claude Fable 5.1: Latent Space reports Astra beats Fable 5.1 on many metrics; Artificial Analysis's independent scores (67 Coding Agent Index, 61 Intelligence Index) place Astra behind Fable 5.1.
  • GPT-6 Astra system card notes improved alignment but decreased chain-of-thought monitorability.
  • Cognition launched SWE-2 (2026-09-10), post-trained from Kimi K3, a 2.8T-parameter base model, reportedly adding 5-6 points on many benchmarks.
  • SWE-2 scores: 50.0% on FrontierCode 1.1 Main (within one point of Fable 5.1 at 64% lower cost), 73.0% on DeepSWE 1.1, and 92.8% on Terminal-Bench 2.1.

Coverage timeline

  1. · 12d ago
    Latent Space· 93
    GPT-6 Astra: an automated AI Engineer you can hire for <$6 an hour

    OpenAI launches GPT-6 Astra, a frontier model scoring 97.6% on FrontierMath and 99.9% on ARC-AGI-3, capable of autonomous AI engineering at roughly $6 per hour.