ZeroHour
Story · 3 sources · 8 articlesfirst updated ()1

GPT-6 Astra launch week: Critical cybersecurity threshold, benchmark leads, Cognition's SWE-2 challenge, and enterprise adoption by Devin and Perplexity

infoModel releaseimportance 92
What's new: Since the previous summary, the story now incorporates the full OpenAI customer case studies published 2026-09-11 and 2026-09-12: Cognition applying Astra to automated testing in Devin (Otter Run simulator recording and coverage report, screenshot-driven bug fixing with verification screenshots, Walden Yan on reduced manual code review) and Perplexity trusting Astra with production systems and…
Merged summary · glm-5.3 · rewritten as coverage arrives

OpenAI launched GPT-6 Astra on 2026-09-09 in ChatGPT Work, Codex, and the API ($10/M input, $50/M output), claiming the first Critical cybersecurity capability threshold, 99.9% on ARC-AGI-3, and the top ErdosBench math score; Cognition countered on 2026-09-10…

OpenAI launched GPT-6 Astra on 2026-09-09 in ChatGPT Work, Codex, and the API, claiming state-of-the-art performance in computer use, browsing, professional work, software engineering, cybersecurity, and science. OpenAI bills Astra as the first model to reach the Critical cybersecurity capability threshold under its Preparedness Framework, with 89% fewer unintended outcomes than GPT-5.6 Sol and 74.7% fewer than Claude Fable 5.1 on its internal computer-use safety benchmark. Pricing is $10 per million input tokens and $50 per million output tokens; OpenAI claims Astra occupies most of the cost-efficiency frontier on Terminal Bench 4.0 and the Artificial Analysis Intelligence Index, and new enterprise admin controls plus plugins from Oracle Analytics, Power BI, Navan, and Avalara shipped alongside. Independent commentary tempered the launch claims: Sebastian Raschka (2026-09-09) called Astra the best model he has used, citing disproportionate gains in 3D rendering, animation, and computer use through the Codex/ChatGPT harness, 99.9% on ARC-AGI-3 versus 7.8% for GPT-5.6 Sol, and leadership on the Artificial Analysis Coding Agent Index, while noting independent aggregate indices show frontier performance without decisive leaps; his article also examined looped transformer/recurrent-depth architecture rumors and speculation that Astra hides its chain-of-thought reasoning. On mathematics, Astra topped ulam.ai's ErdosBench with a score of 3.23, solving 106 of 226 open problems (43 fully solved, 27 disproved) versus GPT-5.6 Sol's 3.12 and 78 solved at maximum reasoning; chief scientist Jakub Pachocki said OpenAI deliberately skipped targeted math optimization to prioritize recursive self-improvement and automated alignment research, benchmark developer Przemek Chojecki estimated gains at 5-10% across tested math-research skills, and Terence Tao warned at the 2026 International Congress of Mathematicians that AI-generated proofs could shift mathematics from proof scarcity to proof overload. Competitive pressure arrived 2026-09-10 when Cognition launched SWE-2, a proprietary 2.8T-parameter mixture-of-experts model with 104B active parameters per token, post-trained from the Kimi K3 base: vendor-reported scores include 50.0% on FrontierCode 1.1 Main, 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1, and 27.3% on Terminal-Bench 4.0, beating Grok 4.6 and SWE-1.7 while claiming to be within one point of Claude Fable 5.1 on FrontierCode at 64% lower cost. Sources…

  • GPT-6 Astra launched 2026-09-09 in ChatGPT Work, Codex, and the API, with claimed state-of-the-art performance in computer use, browsing, professional work, software engineering, cybersecurity, and science
  • Astra is billed as the first model to reach the Critical cybersecurity capability threshold under OpenAI's Preparedness Framework, with 89% fewer unintended outcomes than GPT-5.6 Sol and 74.7% fewer than Claude Fable 5.1 on OpenAI's…
  • Astra pricing: $10 per million input tokens and $50 per million output tokens; OpenAI claims most of the cost-efficiency frontier on Terminal Bench 4.0 and the Artificial Analysis Intelligence Index
  • Enterprise admin controls and plugins from Oracle Analytics, Power BI, Navan, and Avalara launched alongside Astra
  • Sebastian Raschka reports Astra scores 99.9% on ARC-AGI-3 versus 7.8% for GPT-5.6 Sol, leads the Artificial Analysis Coding Agent Index, and shows standout 3D rendering, animation, and computer-use gains, while independent aggregate…
  • Raschka's article analyzes looped transformer/recurrent-depth architecture rumors and speculation that Astra hides its chain-of-thought reasoning
  • Astra tops ulam.ai's ErdosBench with 3.23, solving 106 of 226 open problems (43 fully solved, 27 disproved), versus GPT-5.6 Sol's 3.12 and 78 solved at maximum reasoning
  • Chief scientist Jakub Pachocki said OpenAI deliberately avoided targeted math optimization to prioritize recursive self-improvement and automated alignment research; Przemek Chojecki estimated 5-10% gains across tested math-research skills

Coverage timeline

  1. · 6d ago
    OpenAI News· 92
    GPT-6 Astra: The next generation in intelligence for work

    OpenAI launched GPT-6 Astra, its most capable and aligned model, in ChatGPT Work, Codex, and the API, claiming frontier performance and cybersecurity gains.

  2. · 6d ago
    Hacker News · AI· 92
    GPT-6 Astra, Looped Transformers, and Hidden Reasoning

    OpenAI released GPT-6 Astra, its strongest model to date, with standout 3D rendering and computer-use performance and 99.9% on ARC-AGI-3.

  3. · 5d ago
    The Decoder· 68
    GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design

    OpenAI's GPT-6 Astra tops ulam.ai's ErdosBench math benchmark with 106 of 226 problems solved, while the company prioritizes recursive self-improvement over math optimization.

  4. · 5d ago
    Hacker News · AI· 72
    Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

    Cognition released SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, near Fable 5.1 at 64% lower cost.

  5. · 5d ago
    Hacker News · AI· 62
    Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

    Cognition releases SWE-2, a 2.8T-parameter MoE coding model post-trained from Kimi K3, scoring 92.8 on Terminal-Bench 2.1.

  6. · 4d ago
    OpenAI News· 36
    Cognition helps Devin test its own work with GPT‑6 Astra

    Cognition integrates GPT-6 Astra into Devin, its CLI, and desktop products to automate testing and provide evidence for code review.

  7. · 3d ago
    OpenAI News· 38
    Perplexity trusts GPT-6 Astra with end-to-end systems

    Perplexity uses OpenAI's GPT-6 Astra to craft communications, edit production systems, and generate end-to-end automated tests for its search engine.

  8. · 3d ago
    The Decoder· 32
    GPT-6 Astra needs leaner prompts and fewer guardrails, OpenAI recommends

    OpenAI's Eric Provencher advises developers using GPT-6 Astra to shorten skill descriptions, trim AGENTS.md reading requirements, relax approval rules, and define clear completion goals.