Open-weight AI models near frontier capabilities but lag on safety, report finds
A new report from AI safety nonprofit SaferAI finds that Z.ai's open-weight model GLM-5.2 approaches frontier capabilities in cyber and bio tasks while lacking key safety mitigations, renewing concerns that powerful open models could outpace governance. The model refused none of the offensive tasks it was given. SaferAI's evaluation, run via Z.ai's public API, found GLM-5.2 is only a few months behind leaders like OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7, which refused tasks so consistently that testing could not be completed. Open-weight models allow users to remove safeguards once downloaded, unlike closed systems with API-level controls. Z.ai did not publish a safety framework or risk assessment for GLM-5.2, and did not respond to TechCrunch's questions. China's regulations focus on political content rather than catastrophic risks, while advocates argue open weights aid defense, citing Hugging Face's use of GLM-5.2 against a breach.