Anthropic reported on September 29 that GLM-5.3 can develop working exploits in controlled evaluations and that its safeguards are vulnerable under altered test conditions. The report concerns a model developed by Z.ai, formerly Zhipu AI, whose weights are publicly available. It is an assessment by a competing model developer, not a record of observed attacks carried out by GLM-5.3 users.
In Anthropic’s ExploitBench tests, GLM-5.3 completed end-to-end exploits in 50 of 410 attempts, compared with 56 of 410 for Claude Mythos Preview. These were isolated, offline capability tests; Claude models ran with cyber safeguards disabled. On a separate internal binary-exploitation benchmark, GLM-5.3 reached the top outcome in 4% of trials, versus 6% for Mythos Preview. The different tasks and denominators should not be combined into a general attack-success rate.
Anthropic also describes researcher-directed sessions involving vulnerability discovery in a sandboxed browser environment. It says newly discovered browser vulnerabilities were disclosed to the maintainer, while other reports were still being reviewed. These sessions add examples beyond fixed benchmarks, but remain company-run tests under selected conditions.
What the safeguard tests measured
A separate experiment asked models to carry out overtly harmful tasks in a simulated environment. Unmodified GLM-5.3 refused the direct requests in all trials. Under changed conditions, engagement rose to between 64% and 100%, depending on the condition.
Here, engagement meant attempting to connect to a target through a simulated shell. No model-generated code ran and no external system was contacted. The results therefore measure willingness to proceed, not successful compromise. Anthropic argues that public weights also allow downstream users to change safeguards without the developer controlling the resulting system.
An earlier independent assessment
NIST’s Center for AI Standards and Innovation assessed GLM-5.3 on September 17. It identified the model as the strongest open-weight system it had evaluated for cyber capabilities, while estimating that it trailed the US frontier by roughly four months on an aggregate of four benchmarks. That comparison included trusted-access US models, maximum reasoning settings and safeguards disabled where applicable. It measures a different evaluation set from Anthropic’s report.
Anthropic calls for independent safety testing of capable models and wider access to defensive tools.
