Open-weight AI reached the forefront, but safety did not.
Open-weight AI models have almost matched the forefront in terms of capabilities, but they lag significantly in safety. Once the weights are made public, no research lab can implement safeguards effectively. A recent assessment of China's top open model highlights this disparity.
GLM-5.2, the open-weight model from China's Z.ai, is just a few months behind OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 concerning cyber and bio tasks, as reported by TechCrunch, referencing the safety nonprofit SaferAI. However, it did not decline any of the offensive-cyber or dual-use-biology tasks assigned to it. In contrast, Claude Opus 4.7 refused requests so consistently that SaferAI couldn't even complete the cyber benchmark on that model.
“The frontier of capability is not the frontier of risk,” said SaferAI’s Henry Papadatos, emphasizing that safeguards are just as crucial as the model itself. While Z.ai can secure its hosted services, those protections disappear once someone runs the weights on their own hardware, allowing any safety measures to be removed. This highlights the risks that critics of open models have raised for years.
Closed models also have vulnerabilities. The nonprofit Far.ai discovered numerous universal jailbreaks in xAI’s Grok 4.5 and Google’s Gemini 3.1 Pro. The key difference is that a closed lab can fix a jailbroken model, whereas open weights cannot be recalled once they are released.
To address this, the industry is attempting to implement external safety measures. Recently, Mistral launched Shieldstral, a small open-weight classifier that evaluates text and images against straightforward rules and claims to match the performance of models seven times larger. Similarly, Cisco introduced Antares, open-weight models designed to identify vulnerabilities in code. These open-weight tools are intended to mitigate the risks posed by open weights.
Supporters of openness argue that it can have positive effects. For instance, Hugging Face utilized GLM-5.2 to defend itself during OpenAI’s security breach. Its CEO, Clem Delangue, asserts that systems preventing one type of attack can guard against many more. However, Papadatos believes this is overstated, cautioning that the industry "shouldn't open-source dangerous capabilities,” noting that attackers can adapt quickly while defenders often cannot keep pace.
There is a disconnect in governance frameworks. The White House's new voluntary framework evaluates some closed frontier models for cyber risks, but it reportedly does not currently include open-source models. Anthropic, which previously concentrated on intellectual property theft, has now shifted its focus to safety concerns regarding open weights. Z.ai has yet to release a safety framework for GLM-5.2 and did not respond to inquiries from TechCrunch.
China is aware of the risks but is pursuing a different agenda. Xi Jinping has endorsed open weights while emphasizing the importance of “human control” over AI. As noted by Stanford’s Graham Webster, China’s regulations focus more on political content and social stability than on preventing severe cyber or biological misuse.
The capability competition is nearing resolution: openness is catching up quickly and is much cheaper. However, the safety competition remains unresolved, with AI evolving to both defend and attack. The challenge lies in ensuring that only defensive capabilities are easy to access.
Другие статьи
Open-weight AI reached the forefront, but safety did not.
A report by SaferAI indicated that China's open GLM-5.2 accepted all cyber and biological tasks. The open weights have achieved advanced capabilities, but not necessarily in terms of safety.
