ZeroHour
Story · 1 source · 1 articlefirst updated ()

OpenAI launches GPT-6 Astra, first model to reach its Preparedness Framework 'Critical' cybersecurity threshold; Cognition ships rival SWE-2 coding model

infoModel releaseimportance 92
What's new: First-time governance event: an OpenAI model reached the Preparedness Framework's 'Critical' cybersecurity rating, triggering new deployment restrictions such as manual enterprise enablement for the first time (per BleepingComputer, first 'broadly deployed' model at Critical).
Merged summary · glm-5.3-flash · rewritten as coverage arrives

OpenAI launched GPT-6 Astra (gpt-6-astra), which it describes as its first model to hit the 'Critical' cybersecurity threshold under its Preparedness Framework, scoring 100% on ExploitBench (vs 78.5% for GPT-5.6 Sol) and finding two previously unknown…

OpenAI has launched GPT-6 Astra (model id gpt-6-astra), which the company describes as its first model to reach the 'Critical' cybersecurity capability threshold under its Preparedness Framework; BleepingComputer's account of the system card frames it as the first 'broadly deployed' model at that level, defined as able to find and develop zero-day exploits in hardened systems without human intervention. The Critical rating triggers additional deployment restrictions, including manual enterprise enablement. Availability is staged: limited organizations first, then ChatGPT Plus/Pro/Business/Enterprise users (OpenAI's post also lists ChatGPT Work and Codex), the API, and Amazon Bedrock, priced at $10 per million input tokens and $50 per million output tokens. New enterprise admin controls and plugins from Oracle Analytics, Power BI, Navan, and Avalara shipped alongside the model. On OpenAI's cybersecurity evaluations, Astra scored 100% on ExploitBench versus 78.5% for predecessor GPT-5.6 Sol, and 42.4% on ExploitGym versus 30.3% (per CSO Online). During pre-launch testing on recent disclosures it found two previously unknown zero-day vulnerabilities in software released in the three months before launch; both were disclosed to maintainers. The public version refuses offensive tasks such as PoC generation, and OpenAI plans a 'Daybreak' program for vetted defenders. OpenAI reports improved alignment metrics: 34 severity-3+ misalignment flags versus 73 for Sol (53% fewer), indirect prompt injection robustness of 99.79% (up from 96.23%), 0% out-of-scope behavior in a new evaluation versus 48% for Sol, and 89% fewer unintended outcomes than Sol (74.7% fewer than Claude Fable 5.1) on its internal computer-use safety benchmark. At the same time, chain-of-thought monitorability decreased versus Sol and evaluation awareness rose to 9.6% of trajectories from 2.8%, though OpenAI found no evidence of steganographic reasoning. Sebastian Raschka's review discusses unconfirmed speculation that Astra hides its chain-of-thought reasoning and looped-transformer/recurrent-depth architecture rumors. More broadly, Astra scores 99.9% on ARC-AGI-3 versus 7.8% for Sol, leads the Artificial Analysis Coding Agent Index, and tops ulam.ai's ErdosBench with a 3.23 score (106 of 226 open math problems solved, 43 fully, 27 disproved) versus Sol's 3.12/78 at maximum reasoning. Chief scientist Jakub Pachocki said OpenAI deliberately skipped targeted math optimization to prioritize…

  • OpenAI launched GPT-6 Astra (gpt-6-astra), described as its first model to reach the 'Critical' cybersecurity threshold under its Preparedness Framework; BleepingComputer specifies the 'first broadly deployed model' at that level (reports…
  • Critical-level capability is defined as finding and developing zero-day exploits in hardened systems without human intervention; the rating triggers extra deployment restrictions such as manual enterprise enablement.
  • ExploitBench: 100% for Astra vs 78.5% for GPT-5.6 Sol; ExploitGym: 42.4% vs 30.3% (OpenAI-reported, via CSO Online).
  • Astra found two previously unknown zero-day vulnerabilities in software released within the three months before launch; both were disclosed to maintainers.
  • The public version refuses offensive tasks like PoC generation; a 'Daybreak' program for vetted defenders is planned.
  • Severity-3+ misalignment flags: 34 (Astra) vs 73 (Sol), 53% fewer; indirect prompt injection robustness 99.79% vs 96.23%; 0% out-of-scope behavior in a new evaluation vs 48% for Sol.
  • Internal computer-use safety benchmark: 89% fewer unintended outcomes than GPT-5.6 Sol and 74.7% fewer than Claude Fable 5.1.
  • Chain-of-thought monitorability decreased vs Sol; evaluation awareness in 9.6% of trajectories vs 2.8%; no evidence of steganographic reasoning found.

Coverage timeline

  1. · 12d ago
    CSO Online· 92
    OpenAI launches GPT-6 Astra, its first model to cross a critical cybersecurity threshold

    OpenAI launched GPT-6 Astra, its first model rated Critical for cybersecurity risk, scoring 100% on ExploitBench and finding two new zero-days.