ZeroHour
Story · 1 source · 1 articlefirst updated ()

OpenAI launches GPT-6 Astra, its first model to cross the 'Critical' cybersecurity threshold, scoring 100% on ExploitBench

infoModel releaseimportance 92
What's new: First OpenAI model rated Critical for cybersecurity risk, introducing new deployment restrictions such as manual enterprise enablement. ExploitBench performance rose to 100% from GPT-5.6 Sol's 78.5%, and ExploitGym to 42.4% from 30.3%. First disclosed case of the model finding two previously unknown zero-days (June-August 2026 software) during pre-launch testing. New shipping restriction: the…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

OpenAI released GPT-6 Astra days after it crossed the Critical cybersecurity threshold under the Preparedness Framework; it scored 100% on ExploitBench (vs 78.5% for GPT-5.6 Sol), found two previously unknown zero-days, and ships restricted to defensive…

OpenAI launched GPT-6 Astra, reported September 4, 2026, days after the model reached the Critical cybersecurity capability threshold under its Preparedness Framework — the first model to do so — triggering additional deployment restrictions such as manual enterprise enablement. It scored 100% on ExploitBench (vs 78.5% for predecessor GPT-5.6 Sol), 42.4% on ExploitGym (vs 30.3%), 98% on FrontierMath Tier 4, and 99.9% on ARC-AGI-3, and during pre-launch testing it found two previously unknown zero-day vulnerabilities in software disclosed between June and August 2026 (the three months before launch), demonstrating exploit development against them. The released version is limited to secure code review and patching and refuses PoC exploit requests; unsafeguarded, the model can use unknown vulnerabilities for code execution in hardened browsers and perform OS privilege-escalation exploits. Availability is staged from selected organizations to ChatGPT Plus, Pro, Business and Enterprise, the API (model id gpt-6-astra, $10 per million input / $50 per million output tokens) and Bedrock; The Hacker News lists Azure and AWS Bedrock while CSO Online lists the API, ChatGPT tiers and Amazon Bedrock — a minor discrepancy between sources. OpenAI cites stronger jailbreak robustness, expanded monitoring context and misalignment containment safeguards, but decreased chain-of-thought monitorability versus Sol and 0% out-of-scope behavior in its new evaluation (vs 48% for Sol); safety checks may interrupt legitimate defensive work pending user review. Looser safeguards are planned via Daybreak, including a $1 billion 'Daybreak for Frontline Defenders' program for critical infrastructure and a US MS-ISAC pilot for public sector and water system defenders.

  • GPT-6 Astra is the first OpenAI model to cross the Preparedness Framework's Critical cybersecurity threshold, reached days before the September 4, 2026 launch; it triggers extra deployment restrictions such as manual enterprise enablement…
  • Benchmarks: 100% on ExploitBench vs 78.5% for GPT-5.6 Sol (both reports); 42.4% on ExploitGym vs 30.3% (CSO Online); 98% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3 (The Hacker News).
  • During pre-launch testing the model found two previously unknown zero-day vulnerabilities in software disclosed between June and August 2026 — the three months before launch — and demonstrated exploit development against them (The Hacker…
  • The shipped model is limited to secure code review and patching and refuses offensive tasks such as PoC exploit generation; less restrictive safeguards are planned via OpenAI Daybreak (The Hacker News).
  • Unsafeguarded, the model can use unknown vulnerabilities for code execution in hardened browsers and perform OS privilege-escalation exploits (The Hacker News).
  • API model id gpt-6-astra is priced at $10 per million input tokens and $50 per million output tokens (CSO Online).
  • Rollout: selected organizations first, then ChatGPT Plus, Pro, Business, Enterprise and the API; channel lists differ slightly — The Hacker News adds Azure and AWS Bedrock, while CSO Online lists the API, ChatGPT tiers and Amazon Bedrock.
  • Safety posture: stronger jailbreak robustness, expanded monitoring context and misalignment containment safeguards (The Hacker News); chain-of-thought monitorability decreased vs Sol and out-of-scope behavior measured 0% vs 48% for Sol…

Coverage timeline

  1. · 12d ago
    The Hacker News· 92
    GPT-6 Astra Scores 100% on ExploitBench as OpenAI Blocks PoC Exploit Requests

    OpenAI releases GPT-6 Astra, scoring 100% on ExploitBench, but restricts it to secure code review by blocking PoC exploit generation.