Artificial intelligence

GLM-5.2 matches GPT-5.5 on cyber skills but refuses zero unsafe tasks

2 min read

GLM-5.2 matches GPT-5.5 on cyber skills but refuses zero unsafe tasks
Photo: Logan Voss · Unsplash
0 0
XWhatsAppTelegramLinkedIn

GLM-5.2, a Chinese open-weight AI model from Z.ai, has narrowed the gap with leading frontier systems to just a few months in offensive cyber and biological capabilities — and it refused none of the harmful tasks it was given, according to a report released on August 4, 2026, by AI safety nonprofit SaferAI.

In the same evaluation, Anthropic’s Claude Opus 4.7 refused so consistently that SaferAI could not complete the CyberGym benchmark, a test of cybersecurity skills, on it at all.

The results point to a growing safety gap: while closed models like OpenAI’s GPT-5.5 employ classifiers, refusal training, and API-level controls, such safeguards become unenforceable once open-weight model weights are downloaded and run independently.

“The frontier of capability is not the frontier of risk,” said Henry Papadatos, executive director of SaferAI. “We have to take into account the state of the mitigations to assess the risk properly.”

Z.ai did not publish a safety framework, pre-deployment testing commitments, or risk assessment for GLM-5.2, according to SaferAI. TechCrunch asked Z.ai whether it conducted internal or third-party frontier safety evaluations before release but received no response.

At the World AI Conference in July 2026, Chinese President Xi Jinping emphasized the importance of open-weight models while stressing the need for strict human control over AI.

Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, said Chinese regulations have historically focused on politically sensitive content and social stability rather than catastrophic AI risks. He added that Chinese companies tend to coordinate with regulators behind the scenes, making it difficult to know what internal testing they conduct.

Papadatos stressed that the industry should strive to make only good capabilities easily accessible, noting that attackers adopt new tools faster than defenders. “A ransomware group can change its methods in a week. A hospital cannot,” he said.

Sources

Report / request removal

Related

Comments

No comments yet. Be the first.