ZeroHour
GBHackerspublished ()ingested Mayura Kathir1

Claude Mythos Executes End-to-End Intrusion From Initial Access to Full Domain Compromise

infoModel releaseimportance 75
AI summary · glm-5.3-flash

Anthropic's Claude Mythos Preview, its most cyber-capable model, autonomously completed an end-to-end enterprise intrusion simulation in restricted-access testing.

Anthropic's April 2026 system card describes Claude Mythos Preview as the first model to solve a private cyber range end to end and finish a corporate-network attack simulation an expert would need 10+ hours to complete. It scored 100% pass@1 on a 35-challenge Cybench subset and 0.83 on CyberGym versus 0.67 for Claude Opus 4.6. The model is limited to vetted partners under Project Glasswing; it failed an OT cyber range and could not find novel exploits in a fully patched sandbox.

  • Model chained recon, exploitation, and lateral movement across a weakly defended enterprise range
  • Exploited four distinct bugs in Firefox 147 evaluation versus one for Opus 4.6
  • Restricted release via Project Glasswing for defensive-use cybersecurity partners
  • Anthropic cautions result reflects weak defenses, not hardened-enterprise compromise capability
Full article684 words · extracted from gbhackers.com · click to collapse

Anthropic’s Claude Mythos Preview has demonstrated the ability to complete an end-to-end enterprise intrusion simulation, progressing from initial access through chained exploitation and network traversal to the defined compromise objective.

The result marks a material shift in frontier-model cyber capability: the model did not merely solve isolated CTF-style tasks, but autonomously connected weaknesses commonly found in poorly defended corporate environments.

The finding appears in Anthropic’s April 2026 Claude Mythos Preview System Card, which describes the model as its most cyber-capable release to date.

Anthropic said the model is not being made generally available; instead, it is being provided to a limited group of vetted cybersecurity partners under a defensive-use program called Project Glasswing.

External evaluators tested an early Claude Mythos Preview snapshot across private cyber ranges, capture-the-flag environments, and sandbox-escape scenarios.

Anthropic said Mythos became the first evaluated model to solve one private cyber range end to end and independently completed a corporate-network attack simulation estimated to require an expert operator more than 10 hours to finish.

The enterprise range was designed to reflect weaknesses frequently present in real deployments, including outdated software, security misconfigurations, and reused credentials.

To reach the assigned end-state such as exfiltrating data or disrupting equipment the agent had to identify and execute a sequence of interdependent actions across hosts and network segments.

In practical terms, that means the system could reason across an intrusion chain rather than treating reconnaissance, access, exploitation, movement, and objective completion as disconnected exercises.

End-to-End Intrusion

While the report does not publish a granular step-by-step attack path, the completed simulation supports a high-level scenario in which an agent identifies an exposed weakness, validates access, enumerates reachable systems.

Claude Mythos Preview assisted group achieved a mean score of 4.3 critical failures (Source : GitHub).
Claude Mythos Preview assisted group achieved a mean score of 4.3 critical failures (Source : GitHub).

Such workflows normally require persistent context management, tool use, evidence-based triage, and the ability to abandon unproductive paths.

GitHub Researchers said that, Anthropic said traditional benchmark results are becoming less informative because Claude Mythos Preview saturated nearly all of its CTF-style evaluations.

On a 35-challenge Cybench subset, the model achieved 100% pass@1 across all tested challenges, with 10 trials per challenge.

Its performance was also stronger on CyberGym, a 1,507-task benchmark that measures targeted reproduction of previously disclosed vulnerabilities in real open-source projects.

Claude Mythos Preview scored 0.83 at pass@1, compared with 0.67 for Claude Opus 4.6 and 0.65 for Claude Sonnet 4.6.

The model also showed advanced exploit-development capability in a Firefox 147 evaluation.

Given crash categories and a constrained SpiderMonkey environment, Mythos reliably triaged exploitable bugs and developed proof-of-concept exploits that reached arbitrary code execution.

Synthesis Screening (Source : GitHub).
Synthesis Screening (Source : GitHub).

Anthropic said it successfully leveraged four distinct bugs, while Opus 4.6 could reliably exploit only one.

The report cautions against reading the outcome as evidence that AI can autonomously compromise hardened enterprises.

The successful corporate range represented a small-scale network with a weak security posture, minimal monitoring, no active defenses, and slow response capability.

The range also lacked many defensive layers expected in mature production environments.

Claude Mythos Preview failed to complete a separate operational-technology cyber range and did not discover novel exploits in a properly configured, fully patched sandbox.

These failures underline that sophisticated segmentation, rapid patching, credential hygiene, endpoint telemetry, and active detection still raise meaningful barriers to autonomous intrusion.

For defenders, the principal concern is compression of the attack lifecycle. A capable agent can reduce the time traditionally required to correlate vulnerability data, test attack hypotheses, identify viable routes, and prioritize exploitable conditions.

Security teams should treat exposed legacy services, credential reuse, flat identity permissions, and unmonitored lateral-movement paths as increasingly urgent risks.

Anthropic said it is applying restricted access and probe-based monitoring for prohibited, high-risk dual-use, and broader dual-use activity.

The company expects that general-release models with comparable cyber capabilities would block prohibited activity and, in many cases, high-risk exploit-development requests.

Learn 7 Metric-Gated AI SOC Deployment Phases – Download Free AI SOC Deployment Playbook 2026.

Mayura Kathirhttps://gbhackers.com/

Mayura Kathir is a cybersecurity reporter at GBHackers News, covering daily incidents including data breaches, malware attacks, cybercrime, vulnerabilities, zero-day exploits, and more.

Text extracted automatically; images, tables and formatting may be missing. Original: https://gbhackers.com/end-to-end-intrusion/