Claude Mythos tops first Cyber Weapon Index as only model to complete full cyber kill chain, amid AI outages, OpenAI's $1B defender pledge, and agentic miscues
Booz Allen's first Cyber Weapon Index found Anthropic's Claude Mythos (score 80) was the only one of 18 US and Chinese models to autonomously complete a full cyber kill chain, though all nine frontier API models scored zero on real-world bugs. The same week…
Booz Allen's first Cyber Weapon Index tested 18 US and Chinese AI models as autonomous attackers against a production-grade enterprise network, measuring actions via network and host telemetry and combining vulnerability-research and kill-chain attainment scores. Anthropic's Claude Mythos topped the index at 80 (74 vulnerability research, 86 kill-chain attainment) and was the only model to autonomously complete a full cyber kill chain, moving from a stolen employee credential to administrator-level control in every credentialed attempt and achieving full domain compromise even without credentials; Grok 4.5 (49), GPT 5.6 Sol (46), Muse Spark 1.1 (38), and Kimi K3 (38) followed. Four models reached domain access, four lateral movement, and two credential access; all but one penetrated the network. All nine frontier API models scored zero against real-world bugs versus near-ceiling scores on planted ones; only frontier Anthropic models identified the previously unseen flaw in compiled software, and only Mythos exploited it. Pairing Claude Sonnet with a well-built attack harness rivaled Mythos' performance, suggesting harness quality rivals raw model capability. Booz Allen predicts most tested models will reach Mythos' weaponization level within six months, calls AI-enabled mainstream attacks imminent, and urges sector-specific critical-infrastructure resilience deadlines and US cyber 'overmatch.' The firm stresses this was a controlled benchmark, not evidence of a real-world campaign or victim breach, and advises defenders to enforce least privilege, segmentation, and test containment under service load. Separately, ChatGPT/Codex, Claude, Grok, and Google services suffered rare overlapping outages on Thursday, September 3, 2026. OpenAI cited a routing error for ChatGPT and Codex downtime, but the sources disagree on when it was resolved: WIRED reported unavailability to some users from about 7:43 to 8:17 am PT (~34 minutes; the 7:43 am PT start matches Ars Technica's 10:43 am ET), while Ars Technica reported elevated errors from 10:43 am ET resolved at 12:55 pm ET. Anthropic reported a partial outage with elevated errors on Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 starting 9:23 am ET, identified the cause within about 15 minutes, and resolved it by 12:16 pm ET (9:16 am PT), with a brief Claude Sonnet 5 error spike after noon; Anthropic declined to explain the cause. xAI attributed Grok's outage, starting 6:30 am PT, to an outage at SpaceX's…
- Booz Allen's first Cyber Weapon Index tested 18 US and Chinese AI models as autonomous attackers against a production-grade enterprise network; Claude Mythos scored 80 (74 vulnerability research, 86 kill-chain attainment) and was the only…
- Rankings behind Claude Mythos: Grok 4.5 (49), GPT 5.6 Sol (46), Muse Spark 1.1 (38), Kimi K3 (38); four models reached domain access, four lateral movement, and two credential access, with all but one model penetrating the network.
- All nine frontier API models scored zero against real-world bugs versus near-ceiling scores on planted ones; only frontier Anthropic models identified the previously unseen flaw in compiled software, and only Claude Mythos exploited it.
- Claude Mythos achieved administrator access with stolen credentials in every credentialed attempt and full domain compromise even without credentials; pairing Claude Sonnet with a well-built attack harness rivaled Mythos' performance.
- Booz Allen predicts most tested models will reach Mythos' weaponization level within six months, calls AI-enabled mainstream attacks imminent, and urges sector-specific critical-infrastructure resilience deadlines and cyber overmatch; the…
- On Thursday, September 3, 2026, ChatGPT/Codex, Claude, Grok, and Google services suffered rare overlapping outages; sources agree the ChatGPT/Codex incident began around 10:43 am ET / 7:43 am PT and that OpenAI cited a routing error, but…
- Anthropic's partial outage affected Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 from 9:23 am ET, was identified within about 15 minutes, and was resolved by 12:16 pm ET (9:16 am PT), with a brief Claude Sonnet 5 error spike…
- xAI attributed Grok's outage (starting 6:30 am PT) to an outage at SpaceX's Memphis compute center; DownDetector reports surged from fewer than 10 to 1,365 by 9:45 am; Cloudflare, AWS, and Azure reported no issues and no shared third-party…
Coverage timelineoldest first · each row is one article
- · 13d agoClaude Mythos only model to complete full cyber kill chain, experts say
The Register · Security· 79
Booz Allen's Cyber Weapon Index finds only Claude Mythos completed an autonomous full cyber kill chain; mainstream AI-driven attacks deemed imminent.