GPT-6 Astra launch week: Critical cybersecurity threshold, benchmark leads, Cognition's SWE-2 challenge, and enterprise adoption by Devin and Perplexity
OpenAI launched GPT-6 Astra on 2026-09-09 in ChatGPT Work, Codex, and the API ($10/M input, $50/M output), claiming the first Critical cybersecurity capability threshold, 99.9% on ARC-AGI-3, and the top ErdosBench math score; Cognition countered on 2026-09-10…
OpenAI launched GPT-6 Astra on 2026-09-09 in ChatGPT Work, Codex, and the API, claiming state-of-the-art performance in computer use, browsing, professional work, software engineering, cybersecurity, and science. OpenAI bills Astra as the first model to reach the Critical cybersecurity capability threshold under its Preparedness Framework, with 89% fewer unintended outcomes than GPT-5.6 Sol and 74.7% fewer than Claude Fable 5.1 on its internal computer-use safety benchmark. Pricing is $10 per million input tokens and $50 per million output tokens; OpenAI claims Astra occupies most of the cost-efficiency frontier on Terminal Bench 4.0 and the Artificial Analysis Intelligence Index, and new enterprise admin controls plus plugins from Oracle Analytics, Power BI, Navan, and Avalara shipped alongside. Independent commentary tempered the launch claims: Sebastian Raschka (2026-09-09) called Astra the best model he has used, citing disproportionate gains in 3D rendering, animation, and computer use through the Codex/ChatGPT harness, 99.9% on ARC-AGI-3 versus 7.8% for GPT-5.6 Sol, and leadership on the Artificial Analysis Coding Agent Index, while noting independent aggregate indices show frontier performance without decisive leaps; his article also examined looped transformer/recurrent-depth architecture rumors and speculation that Astra hides its chain-of-thought reasoning. On mathematics, Astra topped ulam.ai's ErdosBench with a score of 3.23, solving 106 of 226 open problems (43 fully solved, 27 disproved) versus GPT-5.6 Sol's 3.12 and 78 solved at maximum reasoning; chief scientist Jakub Pachocki said OpenAI deliberately skipped targeted math optimization to prioritize recursive self-improvement and automated alignment research, benchmark developer Przemek Chojecki estimated gains at 5-10% across tested math-research skills, and Terence Tao warned at the 2026 International Congress of Mathematicians that AI-generated proofs could shift mathematics from proof scarcity to proof overload. Competitive pressure arrived 2026-09-10 when Cognition launched SWE-2, a proprietary 2.8T-parameter mixture-of-experts model with 104B active parameters per token, post-trained from the Kimi K3 base: vendor-reported scores include 50.0% on FrontierCode 1.1 Main, 73.0% on DeepSWE 1.1, 92.8% on Terminal-Bench 2.1, and 27.3% on Terminal-Bench 4.0, beating Grok 4.6 and SWE-1.7 while claiming to be within one point of Claude Fable 5.1 on FrontierCode at 64% lower cost. Sources…
- GPT-6 Astra launched 2026-09-09 in ChatGPT Work, Codex, and the API, with claimed state-of-the-art performance in computer use, browsing, professional work, software engineering, cybersecurity, and science
- Astra is billed as the first model to reach the Critical cybersecurity capability threshold under OpenAI's Preparedness Framework, with 89% fewer unintended outcomes than GPT-5.6 Sol and 74.7% fewer than Claude Fable 5.1 on OpenAI's…
- Astra pricing: $10 per million input tokens and $50 per million output tokens; OpenAI claims most of the cost-efficiency frontier on Terminal Bench 4.0 and the Artificial Analysis Intelligence Index
- Enterprise admin controls and plugins from Oracle Analytics, Power BI, Navan, and Avalara launched alongside Astra
- Sebastian Raschka reports Astra scores 99.9% on ARC-AGI-3 versus 7.8% for GPT-5.6 Sol, leads the Artificial Analysis Coding Agent Index, and shows standout 3D rendering, animation, and computer-use gains, while independent aggregate…
- Raschka's article analyzes looped transformer/recurrent-depth architecture rumors and speculation that Astra hides its chain-of-thought reasoning
- Astra tops ulam.ai's ErdosBench with 3.23, solving 106 of 226 open problems (43 fully solved, 27 disproved), versus GPT-5.6 Sol's 3.12 and 78 solved at maximum reasoning
- Chief scientist Jakub Pachocki said OpenAI deliberately avoided targeted math optimization to prioritize recursive self-improvement and automated alignment research; Przemek Chojecki estimated 5-10% gains across tested math-research skills
Coverage timelineoldest first · each row is one article
- · 6d agoGPT-6 Astra: The next generation in intelligence for work
OpenAI News· 92
OpenAI launched GPT-6 Astra, its most capable and aligned model, in ChatGPT Work, Codex, and the API, claiming frontier performance and cybersecurity gains.
- · 6d agoGPT-6 Astra, Looped Transformers, and Hidden Reasoning
Hacker News · AI· 92
OpenAI released GPT-6 Astra, its strongest model to date, with standout 3D rendering and computer-use performance and 99.9% on ARC-AGI-3.
- · 5d agoGPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design
The Decoder· 68
OpenAI's GPT-6 Astra tops ulam.ai's ErdosBench math benchmark with 106 of 226 problems solved, while the company prioritizes recursive self-improvement over math optimization.
- · 5d agoCognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Hacker News · AI· 72
Cognition released SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, near Fable 5.1 at 64% lower cost.
- · 5d agoCognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
Hacker News · AI· 62
Cognition releases SWE-2, a 2.8T-parameter MoE coding model post-trained from Kimi K3, scoring 92.8 on Terminal-Bench 2.1.
- · 4d agoCognition helps Devin test its own work with GPT‑6 Astra
OpenAI News· 36
Cognition integrates GPT-6 Astra into Devin, its CLI, and desktop products to automate testing and provide evidence for code review.
- · 3d agoPerplexity trusts GPT-6 Astra with end-to-end systems
OpenAI News· 38
Perplexity uses OpenAI's GPT-6 Astra to craft communications, edit production systems, and generate end-to-end automated tests for its search engine.
- · 3d agoGPT-6 Astra needs leaner prompts and fewer guardrails, OpenAI recommends
The Decoder· 32
OpenAI's Eric Provencher advises developers using GPT-6 Astra to shorten skill descriptions, trim AGENTS.md reading requirements, relax approval rules, and define clear completion goals.