ZeroHour
Story · 1 source · 1 articlefirst updated ()

Claude Mythos tops first Cyber Weapon Index as only model to complete full cyber kill chain, amid AI outages, OpenAI's $1B defender pledge, and agentic miscues

mediumAI industryimportance 79
What's new: Previously covered: Cyber Weapon Index results, the September 3 outage timeline, and the start of the agentic business experiment. | Added OpenAI's $1B Daybreak for Frontline Defenders pledge (credits over six months, MS-ISAC pilot, >20x the $50M SLCGP allocation) and the debut of its restricted-capability Astra cybersecurity model. | Added NVIDIA's IFA 2026 announcements: up to 1.9x faster local…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

Booz Allen's first Cyber Weapon Index found Anthropic's Claude Mythos (score 80) was the only one of 18 US and Chinese models to autonomously complete a full cyber kill chain, though all nine frontier API models scored zero on real-world bugs. The same week…

Booz Allen's first Cyber Weapon Index tested 18 US and Chinese AI models as autonomous attackers against a production-grade enterprise network, measuring actions via network and host telemetry and combining vulnerability-research and kill-chain attainment scores. Anthropic's Claude Mythos topped the index at 80 (74 vulnerability research, 86 kill-chain attainment) and was the only model to autonomously complete a full cyber kill chain, moving from a stolen employee credential to administrator-level control in every credentialed attempt and achieving full domain compromise even without credentials; Grok 4.5 (49), GPT 5.6 Sol (46), Muse Spark 1.1 (38), and Kimi K3 (38) followed. Four models reached domain access, four lateral movement, and two credential access; all but one penetrated the network. All nine frontier API models scored zero against real-world bugs versus near-ceiling scores on planted ones; only frontier Anthropic models identified the previously unseen flaw in compiled software, and only Mythos exploited it. Pairing Claude Sonnet with a well-built attack harness rivaled Mythos' performance, suggesting harness quality rivals raw model capability. Booz Allen predicts most tested models will reach Mythos' weaponization level within six months, calls AI-enabled mainstream attacks imminent, and urges sector-specific critical-infrastructure resilience deadlines and US cyber 'overmatch.' The firm stresses this was a controlled benchmark, not evidence of a real-world campaign or victim breach, and advises defenders to enforce least privilege, segmentation, and test containment under service load. Separately, ChatGPT/Codex, Claude, Grok, and Google services suffered rare overlapping outages on Thursday, September 3, 2026. OpenAI cited a routing error for ChatGPT and Codex downtime, but the sources disagree on when it was resolved: WIRED reported unavailability to some users from about 7:43 to 8:17 am PT (~34 minutes; the 7:43 am PT start matches Ars Technica's 10:43 am ET), while Ars Technica reported elevated errors from 10:43 am ET resolved at 12:55 pm ET. Anthropic reported a partial outage with elevated errors on Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 starting 9:23 am ET, identified the cause within about 15 minutes, and resolved it by 12:16 pm ET (9:16 am PT), with a brief Claude Sonnet 5 error spike after noon; Anthropic declined to explain the cause. xAI attributed Grok's outage, starting 6:30 am PT, to an outage at SpaceX's…

  • Booz Allen's first Cyber Weapon Index tested 18 US and Chinese AI models as autonomous attackers against a production-grade enterprise network; Claude Mythos scored 80 (74 vulnerability research, 86 kill-chain attainment) and was the only…
  • Rankings behind Claude Mythos: Grok 4.5 (49), GPT 5.6 Sol (46), Muse Spark 1.1 (38), Kimi K3 (38); four models reached domain access, four lateral movement, and two credential access, with all but one model penetrating the network.
  • All nine frontier API models scored zero against real-world bugs versus near-ceiling scores on planted ones; only frontier Anthropic models identified the previously unseen flaw in compiled software, and only Claude Mythos exploited it.
  • Claude Mythos achieved administrator access with stolen credentials in every credentialed attempt and full domain compromise even without credentials; pairing Claude Sonnet with a well-built attack harness rivaled Mythos' performance.
  • Booz Allen predicts most tested models will reach Mythos' weaponization level within six months, calls AI-enabled mainstream attacks imminent, and urges sector-specific critical-infrastructure resilience deadlines and cyber overmatch; the…
  • On Thursday, September 3, 2026, ChatGPT/Codex, Claude, Grok, and Google services suffered rare overlapping outages; sources agree the ChatGPT/Codex incident began around 10:43 am ET / 7:43 am PT and that OpenAI cited a routing error, but…
  • Anthropic's partial outage affected Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 from 9:23 am ET, was identified within about 15 minutes, and was resolved by 12:16 pm ET (9:16 am PT), with a brief Claude Sonnet 5 error spike…
  • xAI attributed Grok's outage (starting 6:30 am PT) to an outage at SpaceX's Memphis compute center; DownDetector reports surged from fewer than 10 to 1,365 by 9:45 am; Cloudflare, AWS, and Azure reported no issues and no shared third-party…

Coverage timeline

  1. · 13d ago
    The Register · Security· 79
    Claude Mythos only model to complete full cyber kill chain, experts say

    Booz Allen's Cyber Weapon Index finds only Claude Mythos completed an autonomous full cyber kill chain; mainstream AI-driven attacks deemed imminent.