Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits
Anthropic says open-weight GLM-5.3 nearly matches Claude Mythos Preview at building software exploits.
Anthropic reports that Zhipu's downloadable GLM-5.3 nearly matches restricted Claude Mythos Preview on exploit-building benchmarks. On ExploitBench it produced working Chrome V8 exploits in 50 of 410 attempts versus 56, and took full control in 4 percent of an internal OSS-Fuzz binary test versus 6 percent. Anthropic says a paired test found previously unknown browser JavaScript-engine flaws, later reported to the vendor, and that refusal fell sharply under red-team framing and after abliteration. CAISI independently calls GLM-5.3 the most cyber-capable open-weight model so far, roughly four months behind the best US models.
- GLM-5.3 solved 50 of 410 ExploitBench V8 tasks; Mythos Preview solved 56.
- Internal OSS-Fuzz tests: full control in 4 percent versus 6 percent.
- Role-play prompts and abliteration sharply reduced refusal behavior.
- CAISI ranks it the strongest open-weight cyber model, months behind top US models.
Full article852 words · extracted from the-decoder.com · click to collapse
Anthropic deliberately held Mythos Preview back, giving access only to select defenders through Project Glasswing so they could get a head start. The company says those defenders have since found more than 10,000 vulnerabilities in critical software. OpenAI is taking a similar approach with Daybreak. GLM-5.3, on the other hand, is available for anyone to download.
Frontier-level exploits now cost about as much as lunch
ExploitBench measures how well models exploit known bugs in Chrome's V8 engine. On that benchmark, GLM-5.3 built a working exploit in 50 of 410 attempts, while Mythos Preview managed 56. Anthropic also runs an internal binary exploitation benchmark based on open-source projects from Google's OSS-Fuzz. There, GLM-5.3 took full control of the target program in 4 percent of tasks, compared to 6 percent for Mythos Preview. Older models like GLM-5.2 and Claude Opus 4.6 failed both tests, and Kimi K3 and DeepSeek V4.1-Flash barely got off zero.

Anthropic also paired GLM-5.3 with a human expert. Within a single day and with little human attention, the model found several previously unknown vulnerabilities in the JavaScript engine of a widely used browser. It then chained them into a web page that can read any file on a visitor's computer, and in the test it pulled a private SSH key. Anthropic says it reported the vulnerabilities to the browser's developers. Other findings in drivers and device firmware are still under review.
Anthropic used the smaller GLM-5.3-Flash to test how quickly a freshly disclosed vulnerability can be turned into a working attack. The model combined a recently disclosed Chrome bug with another known vulnerability and, with little guidance, built a reliable attack out of them. The attack even got around an extra security feature built into the processor. The whole job took 20 minutes of human attention and eight hours of model time, which would have cost $20.40 at Zhipu's API prices.
The US agency CAISI reached similar conclusions in its own assessment. It calls GLM-5.3 the most cyber-capable open-weight model to date and puts it about four months behind the best US models. That comparison comes with caveats. CAISI tested the US models with their cyber safeguards turned off, and the top tier includes models that only vetted users can access.
Open weights make safeguards easy to remove
In an Anthropic simulation, GLM-5.3 refused openly malicious attack commands. When the same request was dressed up as a red-team exercise, the model tried to connect to the target system in 64 percent of runs. With prefilled reasoning steps, that number rose to 92 percent. After abliteration, a technique that strips refusal behavior out of open weights, it reached 100 percent. The simulation doesn't execute any code, so it can't show whether an attack would actually have succeeded. Protected Claude models stayed at zero.
Anthropic's team says this was its first time using abliteration. The process took about 2,200 GPU hours at a cost of roughly $4,400, and Anthropic estimates an experienced team could do it for around $1,200. The refusal rate for harmful requests fell from over 90 percent to between 2 and 12 percent, while scores on science and cyber tests barely moved. According to Anthropic, several developers had already released unlocked versions within days of the model's launch.
Anthropic concludes that state and non-state actors will likely use models like GLM-5.3 to cause real harm. It points to its own reports and those from other US labs documenting attackers who already use AI. Governments should test capable models, the company argues, and defenders need tools at least as good as the ones their adversaries have.
Anthropic's warning also serves its business
The analysis isn't entirely selfless. Anthropic doesn't release its model weights, and the report presents exactly that as a key security advantage. A cheap Chinese open-weight model that's close to the frontier is also a direct competitor.
The report's conclusions fit Anthropic's business, too. The company wants to bring Claude's cyber capabilities to more defenders, and vetted users already have access to Claude Mythos 5.1. Its call for government testing of GLM-5.3's successors also invites suspicion of regulatory capture, where rules end up mainly protecting established players.
Still, this isn't just self-interest. CAISI's independent assessment backs up the capability numbers, and unlocked versions of the model are already out there.
The UK's AI Security Institute recently found that open models have narrowed their lag in cyber capabilities from six to ten months down to four to seven months. According to the institute, open models are also much cheaper to run, and their safeguards are largely ineffective. AISI warned of a persistent and irreversible misuse risk but also pointed to real benefits like private hosting, customization, and lower costs.
AISI saw that lag as a window for defenders to prepare. At the time, it was still unclear whether open models would also catch up to the leap Mythos Preview represented. The measurements from Anthropic and CAISI suggest GLM-5.3 is a first answer to that question.