Anthropic finds GLM-5.3 builds exploits as safeguards often fail
Anthropic says Zhipu's GLM-5.3 autonomously builds end-to-end exploits, with safeguard bypasses in 64-100% of simulated tests.
Anthropic's analysis of Zhipu AI's open-weight GLM-5.3, which it said was released without meaningful safeguards, finds the model can autonomously develop end-to-end exploits. On ExploitBench, GLM-5.3 succeeded in 50 of 410 attempts, compared with 56 of 410 for Claude Mythos Preview. The reports disagree on the binary control-flow hijack evaluation: one describes a 4% rate on OSS-Fuzz tasks, while a Frontier Red Team account of 100 randomly selected tasks from an internal benchmark gives GLM-5.3 4% and Claude Mythos Preview 6%; both say Claude Opus 4.6 and GLM-5.2 succeeded on none, and the red team said a meaningful cyber-capability threshold has been crossed. Simple techniques bypassed GLM-5.3's safeguards 64-100% of the time in simulated tests, and the weights can be downloaded. NIST's CAISI separately assessed GLM-5.3 as the most cyber-capable open-weight model to date, about four months behind the U.S. frontier. In researcher-driven testing, GLM-5.3 chained multiple discovered zero-days in a popular web browser to steal an SSH private key; one report says that took under a day.
- Zhipu AI's open-weight GLM-5.3 succeeded on 50 of 410 ExploitBench attempts, versus 56 of 410 for Claude Mythos Preview.
- Sources disagree on the binary control-flow hijack test: one cites 4% of OSS-Fuzz tasks for GLM-5.3; Anthropic's Frontier Red Team cites 4% of 100 internal-benchmark trials, with Claude Mythos Preview at 6%.
- Claude Opus 4.6 and GLM-5.2 succeeded on none of those binary-exploitation tasks in both accounts.
- Simple techniques bypassed GLM-5.3's safeguards 64-100% of the time in simulated tests; the weights are downloadable.
- NIST CAISI assessed GLM-5.3 as the most cyber-capable open-weight model to date, about four months behind the U.S. frontier.
- In researcher-driven testing, GLM-5.3 chained multiple discovered browser zero-days to steal an SSH private key; one report says this took under a day.
- Anthropic's Frontier Red Team said a meaningful cyber-capability threshold has been crossed. Reports are dated 2026-09-29.
Coverage timelineoldest first · each row is one article
- · 6h agoGLM-5.3 and the spread of advanced cyber capabilities
Hacker News · AI· 78
Anthropic finds Zhipu's open-weight GLM-5.3 autonomously builds end-to-end exploits and its safeguards can be bypassed 64-100% of the time.
- · 1h agoQuoting Anthropic Frontier Red Team
Simon Willison· 73
Anthropic says GLM-5.3 and Claude Mythos Preview now achieve binary control-flow hijacks.