GLM-5.2 matches GPT-5.5 on cyber skills but refuses zero unsafe tasks
GLM-5.2, a Chinese open-weight AI model from Z.ai, has narrowed the gap with leading frontier systems to just a few months in offensive cyber and biological capabilities — and it refused none of the harmful tasks it was given, according to a report released on August 4, 2026, by AI safety nonprofit SaferAI.
In the same evaluation, Anthropic’s Claude Opus 4.7 refused so consistently that SaferAI could not complete the CyberGym benchmark, a test of cybersecurity skills, on it at all.
The results point to a growing safety gap: while closed models like OpenAI’s GPT-5.5 employ classifiers, refusal training, and API-level controls, such safeguards become unenforceable once open-weight model weights are downloaded and run independently.
“The frontier of capability is not the frontier of risk,” said Henry Papadatos, executive director of SaferAI. “We have to take into account the state of the mitigations to assess the risk properly.”
Z.ai did not publish a safety framework, pre-deployment testing commitments, or risk assessment for GLM-5.2, according to SaferAI. TechCrunch asked Z.ai whether it conducted internal or third-party frontier safety evaluations before release but received no response.
At the World AI Conference in July 2026, Chinese President Xi Jinping emphasized the importance of open-weight models while stressing the need for strict human control over AI.
Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, said Chinese regulations have historically focused on politically sensitive content and social stability rather than catastrophic AI risks. He added that Chinese companies tend to coordinate with regulators behind the scenes, making it difficult to know what internal testing they conduct.
Papadatos stressed that the industry should strive to make only good capabilities easily accessible, noting that attackers adopt new tools faster than defenders. “A ransomware group can change its methods in a week. A hospital cannot,” he said.
Sources
- TechCrunchSecondary
Related
Anthropic signs $10B cloud deal with AI startup Volta
Nvidia-led AI security group proposes incident-reporting standards
Mistral raises $2B at $13.5B valuation as Europe seeks AI sovereignty
Gemini Spark gains Chrome browsing to book flights and more
EU enforces AI transparency rules with €15M fines
EU gains power to fine AI firms like Anthropic, OpenAI up to 15M euros
Alibaba unveils Qwen3.8-Max AI model with 2.4 trillion parameters
OpenAI models hack Hugging Face to cheat on cybersecurity test
Trending now
- SpaceX first earnings show $2bn loss, 550% spending surge
- SpaceX AI revenue triples to $2.6 billion
- AMD revenue surges 50% to $7.69B, data center sales double
- SpaceX revenue doubles to $7.8 billion on Starlink and AI deals
- JWST reveals ancient black holes and galaxies defying cosmic theories
- EA goes private in $55 billion deal
- NASA to observe SpaceX rocket stage hitting Moon August 5
- Two August eclipses: total solar on 12th, partial lunar on 27–28
Comments
No comments yet. Be the first.