ZeroHour
Story · 4 sources · 7 articlesfirst updated ()1

GPT-6 Astra tops ErdosBench and lands in Devin and Perplexity as Cognition ships SWE-2

infoAI industryimportance 72
What's new: No prior merged summary exists; developments newly reported in this window: (1) GPT-6 Astra debuted as the top ErdosBench model at 3.23 (106/226 problems) versus GPT-5.6 Sol's 3.12 (78), with OpenAI stating the lead came without targeted math optimization. (2) Cognition launched SWE-2 on 2026-09-10, claiming first-ever RL scaling to the multi-trillion-parameter regime and near-Fable 5.1…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

OpenAI's GPT-6 Astra leads ulam.ai's ErdosBench math benchmark (106 of 226 problems, score 3.23) without targeted math optimization and is being adopted by Devin and Perplexity, while Cognition launches Kimi K3-based SWE-2, a 2.8T-parameter MoE scoring 50.0…

Reports from 2026-09-10 to 2026-09-12 center on OpenAI's GPT-6 Astra as the new frontier reference point. The Decoder reports Astra tops ulam.ai's ErdosBench with a score of 3.23, solving 106 of 226 open math problems (43 fully solved, 27 disproven), ahead of GPT-5.6 Sol at 3.12 with 78 solved at maximum reasoning; chief scientist Jakub Pachocki says OpenAI deliberately skipped targeted math optimization to prioritize recursive self-improvement and automated alignment research, and Terence Tao warned at the 2026 International Congress of Mathematicians that AI-generated proofs could shift mathematics from proof scarcity to proof overload. On 2026-09-10, Cognition launched SWE-2, a proprietary 2.8T-parameter mixture-of-experts model (104B active per token) post-trained from Kimi K3, with vendor-reported scores of 50.0 on FrontierCode 1.1 Main, 73.0 on DeepSWE 1.1, 92.8 on Terminal-Bench 2.1, and 27.3 on Terminal-Bench 4.0. Sources disagree on SWE-2's standing: one report says it matches Fable 5.1 and GPT-5.6 Sol at a fraction of the price, while another characterizes it as one point behind Fable 5.1 on FrontierCode at a claimed 64% lower cost but trailing Fable 5.1 and GPT-6 Astra by a wide margin on long-horizon Terminal-Bench 4.0 tasks; all figures are vendor-reported and pending independent replication. OpenAI case studies show Devin and Perplexity running Astra for automated testing and production-system edits with less supervision, and OpenAI's Eric Provencher published migration guidance for developers. A Hugging Face paper reports six frontier models converging on one imagined successor architecture under school-audience framing, with a GPT-5.6 Sol output closely overlapping an independently sketched GPT-6 Astra architecture.

  • GPT-6 Astra leads ulam.ai's ErdosBench with a score of 3.23, solving 106 of 226 open math problems (43 fully solved, 27 disproven); GPT-5.6 Sol trails at 3.12 with 78 problems solved at maximum reasoning.
  • OpenAI chief scientist Jakub Pachocki said OpenAI deliberately avoided targeted math optimization for Astra to prioritize recursive self-improvement and automated alignment research; benchmark developer Przemek Chojecki estimated the gain…
  • Terence Tao warned at the 2026 International Congress of Mathematicians that AI-generated proofs could shift mathematics from proof scarcity to proof overload.
  • Cognition's SWE-2 is a proprietary mixture-of-experts model with 2.8T total parameters and 104B active per token, post-trained from the 2.8T-parameter Kimi K3 base; Cognition says post-training added 5-6 points on many benchmarks.
  • Vendor-reported SWE-2 benchmarks: FrontierCode 1.1 Main 50.0, DeepSWE 1.1 73.0, Terminal-Bench 2.1 92.8, Terminal-Bench 4.0 27.3; all figures are pending independent replication.
  • Sources disagree on SWE-2's standing: one report says it matches Fable 5.1 and GPT-5.6 Sol at a fraction of the price and beats Grok 4.6 and SWE-1.7; another says it is one point behind Fable 5.1 on FrontierCode at a claimed 64% lower cost…
  • SWE-2's training scaled reinforcement learning to the multi-trillion-parameter regime for the first time, using Pareto-informed cost penalties that train all reasoning-effort levels in a single run, tripled RL environments, NVFP4/FP8…
  • SWE-2 medium takes 58% fewer turns and costs 81% less than SWE-1.7 on FrontierCode (vendor claim); SWE-2 is available in Devin Desktop and CLI, rolling out on Devin Web and Fusion, with no published weights and no per-token API pricing.

Coverage timeline

  1. · 6d ago
    The Decoder· 68
    GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design

    OpenAI's GPT-6 Astra tops ulam.ai's ErdosBench math benchmark with 106 of 226 problems solved, while the company prioritizes recursive self-improvement over math optimization.

  2. · 6d ago
    Hacker News · AI· 72
    Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

    Cognition released SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, near Fable 5.1 at 64% lower cost.

  3. · 6d ago
    Hacker News · AI· 62
    Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

    Cognition releases SWE-2, a 2.8T-parameter MoE coding model post-trained from Kimi K3, scoring 92.8 on Terminal-Bench 2.1.

  4. · 5d ago
    OpenAI News· 36
    Cognition helps Devin test its own work with GPT‑6 Astra

    Cognition integrates GPT-6 Astra into Devin, its CLI, and desktop products to automate testing and provide evidence for code review.

  5. · 5d ago
    OpenAI News· 38
    Perplexity trusts GPT-6 Astra with end-to-end systems

    Perplexity uses OpenAI's GPT-6 Astra to craft communications, edit production systems, and generate end-to-end automated tests for its search engine.

  6. · 5d ago
    The Decoder· 32
    GPT-6 Astra needs leaner prompts and fewer guardrails, OpenAI recommends

    OpenAI's Eric Provencher advises developers using GPT-6 Astra to shorten skill descriptions, trim AGENTS.md reading requirements, relax approval rules, and define clear completion goals.

  7. · 4d ago
    Hugging Face daily papers· 35
    Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?

    Six frontier models from OpenAI, Anthropic, xAI, and Google DeepMind converge on one imagined successor architecture when asked under a school-audience framing.