Back to News Feed
TechCrunch AI28d agoRebecca Bellan

Open-weight AI models are catching up to the frontier. The safety gap remains.

As global policymakers grapple with the complexities of regulating high-stakes AI systems—such as OpenAI’s GPT-5.6 Sol and Anthropic’s Mythos—a significant shift is occurring in the open-weight landscape. A new analysis from the AI safety nonprofit SaferAI reveals that GLM-5.2, an open-weight model developed by China’s Z.ai, is rapidly closing the performance gap with industry leaders. According to the report, the model trails top-tier systems by only a few months in critical areas like cyber and biological capabilities.

However, this technological convergence has highlighted a widening chasm between raw capability and essential safety infrastructure.

The Safety Disconnect

The SaferAI evaluation, conducted via Z.ai’s public API, paints a concerning picture. When subjected to rigorous testing, GLM-5.2 failed to reject a single request involving offensive cyber operations or dual-use biological tasks. This stands in stark contrast to industry benchmarks; for instance, Claude Opus 4.7 was so robust in its refusal protocols that researchers were unable to complete the CyberGym benchmark—a tool designed to measure cybersecurity proficiency—because the model simply would not engage with the harmful prompts.

This discrepancy serves as a sobering validation for critics who have long argued that the proliferation of open-weight models risks placing potent, unpoliced AI tools into the hands of malicious actors. Once these model weights are downloaded, the ability for developers to enforce safety guardrails effectively vanishes.

"The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly," said Henry Papadatos, executive director of SaferAI.

The Limits of Traditional Safeguards

While closed-model developers like OpenAI and Anthropic rely on a layered defense strategy—including API-level controls, refusal training, and classifiers—these measures are far from infallible. Even with these protections, "jailbreaks" remain a persistent threat.

Recent research from the safety nonprofit Far.ai identified hundreds of "universal jailbreaks" within frontier models such as Google DeepMind’s Gemini 3.1 Pro and xAI’s Grok 4.5. These exploits typically involve complex, multi-step manipulation tactics, such as:

  • Roleplaying to bypass standard behavioral constraints.
  • Authority impersonation to trick the model into compliance.
  • Fabricated conversation histories to confuse the model’s context window.
  • Strategic follow-up prompts that gradually erode safety boundaries.

For open-weight models, these existing defensive strategies are largely irrelevant. Because these models are designed to operate on any hardware, users can strip away safety layers, fine-tune the weights for specific malicious purposes, or alter system prompts to bypass any remaining restrictions.

Navigating the Path to Safer Open AI

The industry is currently debating how to reconcile the benefits of open-source innovation with the inherent risks of powerful AI. Papadatos suggests that the goal should be to ensure that beneficial capabilities remain accessible while systematically stripping away the dangerous ones.

Potential Mitigation Strategies:

  • Pre-training Data Filtering: By scrubbing offensive cybersecurity information from training datasets before the model is even built, developers may be able to reduce hazardous knowledge without significantly degrading the model's overall performance.
  • Selective Restriction: Some developers are opting to limit specific high-risk functions. For example, Anthropic’s Opus 5 is designed to identify vulnerabilities in uncompiled source code but is restricted from analyzing compiled software to prevent offensive exploitation.
  • Rigorous Pre-deployment Testing: Establishing standardized safety frameworks and publishing transparent risk assessments before a model is released to the public.

In the case of GLM-5.2, SaferAI noted a complete absence of a published safety framework or pre-deployment testing commitments. Z.ai did not respond to inquiries regarding whether internal or third-party safety evaluations were conducted prior to the model's release.

The Geopolitical Dimension

The conversation around AI safety is also heavily influenced by regional policy differences. While Chinese leadership has expressed support for open-weight models, they have simultaneously emphasized the need for strict human control.

Graham Webster of the Stanford Cyber Policy Center notes that while China possesses robust AI regulations, their focus has historically been on social stability, misinformation, and political content rather than the "catastrophic" risks associated with cyber or biological warfare.

"U.S. AI thinkers are, in general, more concerned with this existential catastrophic idea than the Chinese community," Webster explained. He noted that many Chinese policy researchers operate under the assumption that if a novel frontier risk emerges, American companies will likely be the first to encounter it. Furthermore, the Chinese regulatory environment relies on real-name attribution and close coordination between companies and the state, which provides a different mechanism for accountability than that found in the West.

The Defense Argument

Proponents of open-weight models, such as Hugging Face CEO Clem Delangue, argue that transparency is a net positive for security. By making powerful models available, organizations can better identify and patch vulnerabilities before attackers exploit them. Hugging Face itself utilized GLM-5.2 to bolster its defenses following a recent security breach.

"The same systems that helped stop an AI-powered cyberattack can now help defend against millions of cyberattacks every day, while helping us identify and fix vulnerabilities before attackers exploit them," Delangue noted in a recent social media statement.

However, Papadatos remains skeptical that this justifies the widespread release of dangerous capabilities. He warns that the speed at which attackers adopt new tools often outpaces the defensive capabilities of organizations. "The main point in my mind is that we shouldn’t just accept that dangerous capabilities are easily accessible by anyone anywhere," he concluded. "We shouldn't just accept that dangerous capabilities are easily accessible by anyone anywhere."

As the gap between frontier performance and safety mitigations continues to narrow, the debate is shifting from whether these models can compete with the best in the world to how society can possibly manage the risks once the genie is out of the bottle.

#model