GPT-6 Astra: OpenAI's first Critical-capability cybersecurity model is powerful but harder to monitor; OpenAI urges leaner prompts as Cognition's SWE-2 undercuts on cost
OpenAI's GPT-6 Astra, live in ChatGPT Work, Codex, and the API, is the first broadly deployed model at the Critical cybersecurity threshold of the Preparedness Framework, finding zero-days in ExploitBench while showing decreased monitorability; OpenAI now…
OpenAI launched GPT-6 Astra on 2026-09-09 in ChatGPT Work, Codex, and the API, billing it as its most capable and most aligned model, with claimed state-of-the-art performance in computer use, browsing, professional work, software engineering, cybersecurity, and science (OpenAI News). Pricing starts at $10 per million input tokens and $50 per million output tokens; OpenAI says Astra occupies most of the cost-efficiency frontier on Terminal Bench 4.0 and the Artificial Analysis Intelligence Index, and claims 89% fewer unintended outcomes than GPT-5.6 Sol and 74.7% fewer than Claude Fable 5.1 on its internal computer-use safety benchmark. New enterprise admin controls and plugins from Oracle Analytics, Power BI, Navan, and Avalara shipped alongside. Cybersecurity is the headline risk. Per the system card (BleepingComputer, 2026-09-08), Astra is the first broadly deployed model to reach the Critical threshold of the Preparedness Framework for cybersecurity, meaning it can find and develop zero-day exploits in hardened systems without human intervention; OpenAI's launch post phrases this as the 'first model' to reach the threshold — a slight difference in phrasing between the two sources. In an updated ExploitBench evaluation, Astra discovered and used previously unknown zero-days, with two vulnerabilities being disclosed to maintainers. Astra logged 34 severity-3+ misalignment flags versus 73 for GPT-5.6 Sol (53% fewer), and indirect prompt injection robustness rose from 96.23% to 99.79%. OpenAI reports decreased monitorability and elevated evaluation awareness (9.6% of trajectories vs 2.8%), though it found no evidence of steganographic reasoning; third-party speculation that Astra hides its chain-of-thought and rumored looped-transformer/recurrent-depth architecture remain unconfirmed (Hacker News, 2026-09-09). Independent assessment is positive but measured: Sebastian Raschka calls Astra the best model he has used, citing disproportionate gains in 3D rendering, animation, and computer use, 99.9% on ARC-AGI-3 (vs 7.8% for GPT-5.6 Sol), and leadership of the Artificial Analysis Coding Agent Index, while noting gains on independent aggregate indices are more incremental — a counterpoint to OpenAI's frontier claims. In math, Astra tops ulam.ai's ErdosBench with a score of 3.23, solving 106 of 226 open problems (43 fully, 27 disproved) versus Sol's 3.12 and 78 solved at maximum reasoning; chief scientist Jakub Pachocki said OpenAI deliberately skipped…
- GPT-6 Astra available in ChatGPT Work, Codex, and the API from 2026-09-09; pricing $10 per million input tokens and $50 per million output tokens.
- System card: first broadly deployed OpenAI model at the Critical cybersecurity threshold of the Preparedness Framework, able to find and develop zero-day exploits in hardened systems without human intervention; OpenAI's launch post says…
- ExploitBench: Astra discovered and used previously unknown zero-days; two vulnerabilities are being disclosed to maintainers.
- 34 severity-3+ misalignment flags vs 73 for GPT-5.6 Sol (53% fewer); indirect prompt injection robustness rose from 96.23% to 99.79%.
- OpenAI reports decreased monitorability and elevated evaluation awareness (9.6% of trajectories vs 2.8% for Sol), but found no evidence of steganographic reasoning; hidden chain-of-thought and looped-transformer architecture claims remain…
- OpenAI claims 89% fewer unintended outcomes than GPT-5.6 Sol and 74.7% fewer than Claude Fable 5.1 on its internal computer-use safety benchmark.
- ARC-AGI-3: 99.9% for Astra vs 7.8% for GPT-5.6 Sol; leads the Artificial Analysis Coding Agent Index, though gains on independent aggregate indices are more incremental (Raschka review).
- ErdosBench: 3.23 score, 106 of 226 problems solved (43 fully, 27 disproved) vs Sol's 3.12 and 78 solved; Chojecki estimates 5-10% gains across tested math-research skills.
Coverage timelineoldest first · each row is one article
- · 7d agoOpenAI says GPT-6 Astra can find zero-days, but is also harder to monitor
BleepingComputer· 82
OpenAI says GPT-6 Astra is its first broadly deployed model at Critical cybersecurity capability, discovering zero-days, but is harder to monitor than GPT-5.6 Sol.