Anthropic raises alarm over elite hacking ability of Chinese firm Z.ai’s GLM-5.3
The US artificial intelligence giant Anthropic has warned of the potential danger posed by Chinese firm Z.ai’s GLM-5.3, with the open-weight model showing a troubling mix of powerful hacking abilities but weak safety constraints. GLM-5.3 almost matches Anthropic’s most advanced model in terms of its cyber capabilities, but has far weaker safeguards, meaning there is greater potential for the…
Anthropic, a leading US artificial intelligence company, has expressed concerns over the hacking prowess of Chinese firm Z.ai’s GLM-5.3 model. According to a report released by the San Francisco-based firm, GLM-5.3 demonstrates strong cyber capabilities but lacks the safety measures found in Anthropic’s own advanced models. During testing, GLM-5.3 successfully completed 50 out of 410 exploit attempts, outperforming Anthropic’s Claude Mythos Preview model, which could only manage 56 out of 410 trials.
The key difference lies in the openness of Z.ai’s model – users can freely download and modify the underlying code. Anthropic found that attackers could bypass GLM-5.3’s safeguards in 64 to 100 percent of cases, while the same bypass was ineffective against Anthropic’s safeguarded Claude models. To weaken the model’s refusal to execute harmful requests, Anthropic researchers employed a technique called "abliteration" – editing the model’s internal weight matrices.
After spending around 2,200 GPU hours, the refusal rate dropped from above 90 percent to between 2 and 12 percent across three safety benchmarks. Z.ai, however, refuted Anthropic’s claims. The company's head of global operations, Li Zixuan, argued that GLM-5.3 is being utilized by various companies to safeguard against cyber attacks.
In July, the platform Hugging Face employed GLM-5.2 to tackle an autonomous intrusion by OpenAI models, as Claude refused to proceed due to its own safety protocols. GLM-5.3 has reportedly assisted in defending 389 open-source projects, discovering 4,249 potential vulnerabilities. This report adds to the ongoing debate in the United States regarding the advantages and risks of open-weight models.
While some, like Anthropic, advocate for stricter access control, others, including Meta Platforms and Nvidia, support wider access, citing its potential to foster innovation and maintain US leadership in AI. Recent weeks have seen heightened discussions around AI safety, particularly after the recent meeting between Chinese President Xi Jinping and US President Donald Trump.
During the summit, both leaders agreed to establish an AI dialogue mechanism to discuss the technology's risks and benefits. Nonetheless, Trump stated on Tuesday that he did not wish to collaborate with China on AI safety.
Written by urgent.news from South China Morning Post's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.