Anthropic says Chinese AI model combines strong hacking skills with weak safeguards

Anthropic says Z.ai’s open-weight GLM-5.3 nearly matched one of its advanced models in tests of end-to-end cyber exploits while remaining easier to bypass. Z.ai rejected the warning, highlighting the model’s reported use in defending open-source projects against cyber threats.
Anthropic has warned that Chinese company Z.ai’s GLM-5.3 combines advanced cyber capabilities with safeguards that are comparatively easy to bypass. In a report released on Tuesday, the US AI company said it tested the open-weight model’s ability to build end-to-end cyber exploits. GLM-5.3 successfully completed 50 of 410 exploit attempts, compared with 56 completed by Anthropic’s Claude Mythos Preview.
Anthropic said Mythos is available only to vetted users, while GLM-5.3 can be downloaded and modified because its underlying model weights are open. The company said simple techniques bypassed GLM-5.3’s safeguards between 64% and 100% of the time, whereas the same attacks did not succeed against safeguarded Claude models. Anthropic researchers also used a technique known as “abliteration” to remove the model’s refusal mechanisms.
After about 2,200 GPU hours of work, the model’s refusal rate fell from above 90% to between 2% and 12% across three safety benchmarks. Anthropic estimated that the effort cost about $4,400, while an experienced team might accomplish it for roughly $1,200. Z.ai’s head of global operations, Li Zixuan, rejected the implication that the model was only a threat.
He said the company’s models were being used by businesses to defend against cyber attacks and that GLM-5.3 had helped protect 389 open-source projects and identify 4,249 potential vulnerabilities. The report comes amid a wider US debate about open-weight models. Anthropic is advocating greater control over frontier AI access, while Meta Platforms and Nvidia have supported open models as a way to accelerate innovation.
The Centre for AI Standards and Innovation recently described GLM-5.3 as the most cyber-capable open-weight model released to date, while estimating that it remained about four months behind the overall US frontier on combined cybersecurity benchmarks.
This independently written report is based on information supplied by the named publisher. Vertrix News has not independently verified the source report.