by Vivek Gupta - 16 hours ago - 5 min read
Open-weight artificial intelligence is moving much faster than many researchers and policymakers expected. Models that developers can download, modify and run independently are no longer trailing systems from OpenAI, Anthropic and Google by a year or more. In several technically sensitive areas, the gap has narrowed to only a few months.
A new evaluation of China-based Z.ai’s GLM-5.2 model illustrates how quickly this shift is happening. The model demonstrated cyber and biological capabilities approaching those of recent frontier systems, but it did so without many of the access controls, refusal mechanisms and monitoring systems used by closed-model providers.
The result is an increasingly uncomfortable imbalance: advanced AI capabilities are becoming widely accessible faster than reliable safety protections are being developed.
GLM-5.2 was released by Z.ai, formerly known as Zhipu AI, on June 16, 2026. SaferAI independently evaluated the model across cyber offence, biological risk, loss of control and harmful manipulation, the four systemic risk areas covered by the European Union’s General-Purpose AI Code of Practice.
According to SaferAI, GLM-5.2 was generally only two to four months behind the leading closed models available when it launched. Its biological knowledge was approximately level with Anthropic’s Claude Opus 4.7 and slightly below OpenAI’s GPT-5.5. Its cyber capabilities were comparable to Claude Opus 4.6 and close to GPT-5.5, although its software-engineering performance remained further behind.
Other evaluations point in the same direction. The UK AI Security Institute found that GLM-5.2 performed similarly to closed models released four to seven months earlier. That is considerably narrower than the six-to-ten-month lag the institute measured for leading open-weight models during much of 2025.
Epoch AI reached a similar conclusion using its broader Epoch Capabilities Index. Since January 2026, the strongest open-weight models have trailed the closed frontier by an average of four months, equivalent to an eight-point difference on its aggregate capability index.
Cybersecurity produced some of the report’s most significant findings. GLM-5.2 performed near the saturation point on Cybench, a benchmark covering areas such as reverse engineering, exploitation and web security.
On the more demanding CyberGym evaluation, its reproduction rate increased from 36.6% with a two-million-token budget to 76.2% when the budget was raised to 50 million tokens. The improvement suggests that giving an AI agent more time and computing resources can substantially increase its ability to complete complex cyber tasks.
SaferAI also found that GLM-5.2 did not refuse the offensive-security tasks included in its evaluation. Claude Opus 4.7, by comparison, refused so consistently that researchers were unable to complete the CyberGym evaluation using that model.
NIST’s Center for AI Standards and Innovation separately concluded that GLM-5.2’s safeguards allowed assistance with agentic cyber exploit development. The agency found that its overall capabilities were similar to OpenAI’s GPT-5.2, while its cyber performance resembled Anthropic’s Opus 4.6.
These results do not prove that GLM-5.2 can autonomously execute sophisticated real-world attacks. Public benchmarks cover only selected capabilities and may not reflect the complexity of live systems. They do, however, suggest that advanced offensive knowledge is becoming available in models that can eventually be operated beyond the control of their original developer.
The biological evaluations showed a similarly narrow gap. GLM-5.2 met or exceeded the PhD-level human-expert baseline across every LAB-Bench subtask tested by SaferAI.
On BioMysteryBench, it solved approximately 81% of the human-solvable problems it completed and around one-third of the problems that no participating human expert had solved. Its overall performance was close to Claude Opus 4.7 and GPT-5.5, both released roughly two months earlier.
These benchmarks mainly test scientific reasoning rather than the complete process required to create a biological threat. They do not measure whether a model can acquire materials, operate laboratory equipment or successfully complete an end-to-end biological programme.
The concern is nevertheless growing because open-weight systems can be fine-tuned, stripped of safeguards and combined with external tools without provider oversight.
Closed AI platforms can monitor requests, apply content filters, restrict accounts, impose rate limits and update a model when a vulnerability is discovered. Those protections are imperfect, but providers retain the ability to intervene.
Open-weight models function differently. Once the weights have been downloaded, users can modify the model, remove refusal training or operate it on private infrastructure. The original developer cannot reliably monitor its use or recall every distributed copy.
The International AI Safety Report 2026 describes this release process as effectively irreversible. It also notes that open-weight systems can expand research access, improve privacy and allow smaller developers to adapt advanced models for local languages and specialised applications. The same flexibility, however, makes safeguards easier to remove and malicious use more difficult to trace.
This does not mean open-weight AI is inherently unsafe. Independent researchers can inspect models, discover vulnerabilities and build defensive tools that would be impossible with completely closed systems. The risk depends on the model’s capability, the strength of its built-in protections and the additional danger created by releasing its weights.
The findings arrive as governments debate whether open-weight systems should face the same evaluations as closed frontier models.
The Trump administration recently told AI companies that open-weight models would not be included in its planned voluntary federal safety-testing framework, according to Reuters. Critics argue that excluding downloadable models could create a major oversight gap precisely as their capabilities approach the frontier.
The broader safety environment is already under pressure. Stanford’s 2026 AI Index recorded 362 documented AI incidents, up from 233 in 2024, while responsible-AI reporting remained inconsistent among major developers.
SaferAI stopped short of declaring GLM-5.2 dangerous, stressing that its results cover only a subset of possible risks. Its evaluation nevertheless supports a larger conclusion: measuring capability alone is no longer enough.
As open-weight models move within months of the frontier, safety testing, release documentation and misuse evaluations will need to happen before distribution, not after powerful weights have already spread across the internet.