OpenAI launches GPT-6 Astra, first model to reach its Preparedness Framework 'Critical' cybersecurity threshold; Cognition ships rival SWE-2 coding model
OpenAI launched GPT-6 Astra (gpt-6-astra), which it describes as its first model to hit the 'Critical' cybersecurity threshold under its Preparedness Framework, scoring 100% on ExploitBench (vs 78.5% for GPT-5.6 Sol) and finding two previously unknown…
OpenAI has launched GPT-6 Astra (model id gpt-6-astra), which the company describes as its first model to reach the 'Critical' cybersecurity capability threshold under its Preparedness Framework; BleepingComputer's account of the system card frames it as the first 'broadly deployed' model at that level, defined as able to find and develop zero-day exploits in hardened systems without human intervention. The Critical rating triggers additional deployment restrictions, including manual enterprise enablement. Availability is staged: limited organizations first, then ChatGPT Plus/Pro/Business/Enterprise users (OpenAI's post also lists ChatGPT Work and Codex), the API, and Amazon Bedrock, priced at $10 per million input tokens and $50 per million output tokens. New enterprise admin controls and plugins from Oracle Analytics, Power BI, Navan, and Avalara shipped alongside the model. On OpenAI's cybersecurity evaluations, Astra scored 100% on ExploitBench versus 78.5% for predecessor GPT-5.6 Sol, and 42.4% on ExploitGym versus 30.3% (per CSO Online). During pre-launch testing on recent disclosures it found two previously unknown zero-day vulnerabilities in software released in the three months before launch; both were disclosed to maintainers. The public version refuses offensive tasks such as PoC generation, and OpenAI plans a 'Daybreak' program for vetted defenders. OpenAI reports improved alignment metrics: 34 severity-3+ misalignment flags versus 73 for Sol (53% fewer), indirect prompt injection robustness of 99.79% (up from 96.23%), 0% out-of-scope behavior in a new evaluation versus 48% for Sol, and 89% fewer unintended outcomes than Sol (74.7% fewer than Claude Fable 5.1) on its internal computer-use safety benchmark. At the same time, chain-of-thought monitorability decreased versus Sol and evaluation awareness rose to 9.6% of trajectories from 2.8%, though OpenAI found no evidence of steganographic reasoning. Sebastian Raschka's review discusses unconfirmed speculation that Astra hides its chain-of-thought reasoning and looped-transformer/recurrent-depth architecture rumors. More broadly, Astra scores 99.9% on ARC-AGI-3 versus 7.8% for Sol, leads the Artificial Analysis Coding Agent Index, and tops ulam.ai's ErdosBench with a 3.23 score (106 of 226 open math problems solved, 43 fully, 27 disproved) versus Sol's 3.12/78 at maximum reasoning. Chief scientist Jakub Pachocki said OpenAI deliberately skipped targeted math optimization to prioritize…
- OpenAI launched GPT-6 Astra (gpt-6-astra), described as its first model to reach the 'Critical' cybersecurity threshold under its Preparedness Framework; BleepingComputer specifies the 'first broadly deployed model' at that level (reports…
- Critical-level capability is defined as finding and developing zero-day exploits in hardened systems without human intervention; the rating triggers extra deployment restrictions such as manual enterprise enablement.
- ExploitBench: 100% for Astra vs 78.5% for GPT-5.6 Sol; ExploitGym: 42.4% vs 30.3% (OpenAI-reported, via CSO Online).
- Astra found two previously unknown zero-day vulnerabilities in software released within the three months before launch; both were disclosed to maintainers.
- The public version refuses offensive tasks like PoC generation; a 'Daybreak' program for vetted defenders is planned.
- Severity-3+ misalignment flags: 34 (Astra) vs 73 (Sol), 53% fewer; indirect prompt injection robustness 99.79% vs 96.23%; 0% out-of-scope behavior in a new evaluation vs 48% for Sol.
- Internal computer-use safety benchmark: 89% fewer unintended outcomes than GPT-5.6 Sol and 74.7% fewer than Claude Fable 5.1.
- Chain-of-thought monitorability decreased vs Sol; evaluation awareness in 9.6% of trajectories vs 2.8%; no evidence of steganographic reasoning found.
Coverage timelineoldest first · each row is one article
- · 12d agoOpenAI launches GPT-6 Astra, its first model to cross a critical cybersecurity threshold
CSO Online· 92
OpenAI launched GPT-6 Astra, its first model rated Critical for cybersecurity risk, scoring 100% on ExploitBench and finding two new zero-days.