Quoting Anthropic Frontier Red Team
Anthropic says GLM-5.3 and Claude Mythos Preview now achieve binary control-flow hijacks.
Anthropic’s Frontier Red Team evaluated several models on 100 randomly selected tasks from an internal binary-exploitation benchmark. GLM-5.3 developed full control-flow hijacks in 4% of trials, and Claude Mythos Preview did so in 6%. Earlier models Claude Opus 4.6 and GLM-5.2 succeeded on none of the tasks. The team said a meaningful cyber-capability threshold has been crossed.
- One hundred tasks were drawn from an internal binary-exploitation benchmark.
- GLM-5.3 achieved full control-flow hijacks in 4% of trials.
- Claude Mythos Preview succeeded in 6% of trials.
- Claude Opus 4.6 and GLM-5.2 succeeded on none of the tasks.
We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. — Anthropic Frontier Red Team , GLM-5.3 and the spread of advanced cyber capabilities Tags: anthropic , generative-ai , ai-security-research , glm , ai , ai-in-china , llms
This source does not provide full text. Read it at simonwillison.net.