Anthropic says GLM-5.3, an open-weight model from Zhipu AI, also known as Z.ai, has reached a level of cyber capability comparable to some of its restricted frontier models. In a September 29 research report, the company argues that the model's availability and weaknesses in its safeguards change which capabilities potential attackers can access.
The report credits Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao and Tripp Gallagher. Their account combines automated evaluations with researcher-guided experiments conducted in isolated, sandboxed environments against offline targets. The findings are Anthropic's, not an independent Franklin AI reproduction or a report of attacks on unsuspecting users.
Exploit-development tests show a capability increase
On ExploitBench, Anthropic reports that GLM-5.3 completed end-to-end exploits in 50 of 410 attempts, compared with 56 of 410 for Claude Mythos Preview. The evaluation uses known vulnerabilities in the V8 engine associated with Google Chrome. Those counts concern success within the benchmark's tested attempts, rather than a general rate for compromising browsers in the wild.
A separate internal binary-exploitation evaluation covered 100 randomly selected tasks involving open-source projects. Anthropic reports full control-flow hijacks in 4% of GLM-5.3 trials and 6% for Mythos Preview. It says earlier models, including GLM-5.2 and Claude Opus 4.6, did not achieve that outcome in the evaluation.
The capability charts used Claude models with safeguards disabled. That condition is essential to interpreting the comparison. A model's ability when restrictions are removed is different from what an ordinary user can obtain through its safeguarded service.
Researcher-guided examples have narrower boundaries
Anthropic also describes experts using the models to find and develop exploits against targets where the experts did not know of existing vulnerabilities. The sessions ran in controlled environments, with limited human attention but more elapsed model time.
One example concerns previously unknown flaws in a local Linux browser build. The researchers say they disclosed those vulnerabilities to the maintainer. Their belief that other platforms might also be affected is a possibility in the report, not a demonstrated exploit against each platform.
Another example uses the smaller GLM-5.3-Flash model and already disclosed flaws. The company reports 20 minutes of human attention alongside eight hours of model work, with an estimated $20.40 charge at Zhipu's API prices. That is a reported session cost under specified conditions, not a current fixed price for an attack or a claim that the human-only time measures the whole process.
These examples add context to the automated scores. They do not establish a universal capability, and they should remain distinct from evidence of real-world malicious use.
Safeguard bypass results measure engagement, not completed attacks
Anthropic says unmodified GLM-5.3 refused overtly malicious requests in every trial of its simulated-world test. Under the different bypass conditions it tested, the model instead engaged between 64% and 100% of the time.
Those percentages measure whether the model tried to connect to a remote target in the simulated setting. They are not rates of successful compromise. The figure description specifies 50 samples per condition, built from five attack orders, two targets and five attempts.
Anthropic reports that the same tested approaches did not cause safeguarded Claude models to carry out the harmful tasks. Some approaches also depend on access that the Claude API does not provide, such as modifying model weights. The comparison therefore concerns both model behavior and differences in how the systems are released.
The report argues that open weights allow users to change or remove refusal behavior while retaining much of the model's capability. This article summarizes that finding without reproducing the operational bypass procedures. The result raises a deployment concern, but it does not prove that hosted safeguards are impossible to defeat or that every open-weight model shares the same weakness.
Access conditions are part of the comparison
Anthropic cites a September 17 assessment from NIST's Center for AI Standards and Innovation, saying its broad capability findings agree with that assessment. Its account says CAISI compared against some US models with cyber safeguards disabled and included models available only to vetted users. That is Anthropic's description of CAISI's work; the cited assessment was not separately evaluated here.
Anthropic's policy argument is that broadly available cyber capability needs stronger safeguards and independent safety testing. It also says defenders can benefit from the same capabilities and calls for wider access to advanced defensive models through trusted programs.
The report supplies controlled results and a clear argument from a competing model provider. Independent evaluation remains important, especially when comparing a downloadable model with restricted services. Capability scores, safeguards and actual access each answer a different question; readers should keep all three in view before treating one benchmark comparison as a complete account of cyber risk.
