Anthropic warns Chinese open-weight model GLM-5.3 has frontier-level hacking ability
Anthropic says Z.ai's open-weight GLM-5.3 nearly matches its vetted-only Claude Mythos Preview at building exploits, with safeguards that can be bypassed up to 100% of the time.
Anthropic warned on Tuesday 29 September that Chinese firm Z.ai’s (Zhipu AI’s) open-weight model GLM-5.3, released free in late August, has frontier-level cyberattack capabilities with weak safeguards (https://finance.biggo.com/news/fadcd07a-e46e-45f3-8029-b7f132f5e7d5).
On Anthropic’s ExploitBench test of building exploits for known Chrome V8 vulnerabilities, GLM-5.3 succeeded in 50 of 410 attempts, close to the 56 of Claude Mythos Preview — Anthropic’s vetted-only frontier model. Claude Opus 4.6, GLM-5.2, Kimi K3 and DeepSeek-V4.1-Flash scored near zero (https://the-decoder.com/anthropic-says-zhipus-open-weight-glm-5-3-nearly-matches-claude-mythos-preview-at-building-exploits/).
Anthropic found GLM-5.3’s safeguards bypassed 64% of the time when attack instructions were dressed up as red-team exercises, 92% with pre-filled reasoning, and 100% after “abliteration” — editing internal weight matrices to strip refusal behaviour. The same techniques failed against safeguarded Claude models. The abliteration took about 2,200 GPU hours, costing roughly $4,400 (about $1,200 for an experienced team), and cut refusal rates from above 90% to 2–12% across JailbreakBench, HarmBench and StrongREJECT, with barely any capability loss (https://www.scmp.com/tech/big-tech/article/3369354/anthropic-raises-alarm-over-elite-hacking-ability-chinese-firm-zais-glm-53).
In a demo, the smaller GLM-5.3-Flash built a working exploit chain for an ARM64 target from public details of CVE-2026-11645, with about 20 minutes of human attention, eight hours of model time and $20.40 in Zhipu API fees, per BigGo Finance (https://finance.biggo.com/news/fadcd07a-e46e-45f3-8029-b7f132f5e7d5).
The US Center for AI Standards and Innovation said on 17 September that GLM-5.3 was the most cyber-capable open-weight model to date, about four months behind the best US models — US models tested with safeguards off, vetted-only models included (https://the-decoder.com/anthropic-says-zhipus-open-weight-glm-5-3-nearly-matches-claude-mythos-preview-at-building-exploits/). Anthropic notes its tests were simulated and do not fully reproduce real-world behaviour.
Sources
- South China Morning Post: https://www.scmp.com/tech/big-tech/article/3369354/anthropic-raises-alarm-over-elite-hacking-ability-chinese-firm-zais-glm-53
- The Decoder: https://the-decoder.com/anthropic-says-zhipus-open-weight-glm-5-3-nearly-matches-claude-mythos-preview-at-building-exploits/
- BigGo Finance: https://finance.biggo.com/news/fadcd07a-e46e-45f3-8029-b7f132f5e7d5
More on this topic: all Technology stories