An Open-Source Chinese Model Just Claimed Parity on Finding Software Vulnerabilities
Z.ai says GLM-5.3 scored 84.5% on CyberGym against 83.8% for Anthropic's restricted Mythos 5 — and plans to ship the weights.
An Open-Source Chinese Model Just Claimed Parity on Finding Software Vulnerabilities
Chinese AI startup Z.ai says its open-source GLM-5.3 scored 84.5% on CyberGym, against 83.8% for Anthropic's access-restricted Mythos 5.
CyberGym tests whether a model can review code, identify security flaws, and confirm those flaws are real. The confirmation step matters — it separates a scanner that generates noise from one that generates work worth doing.
The results have not been independently verified, and that caveat should carry through everything below.
Why the 0.7-point gap is not the point
The headline is parity. The actual news is the distribution model.
Mythos 5 is restricted, with initial access limited to a small set of partners — a safety decision, and a defensible one, since vulnerability discovery at frontier capability is close to a textbook dual-use case. GLM-5.3 is open-source, and Z.ai plans a public release after a two-week safety review.
If the benchmark holds up, the capability one lab decided to gate becomes downloadable. Not through a leak or a jailbreak, but through a different lab making a different call.
The dual-use problem in its purest form
Vulnerability discovery is genuinely symmetric. The same capability that lets a defender audit a codebase before shipping lets an attacker audit it afterwards.
The usual argument for open weights — that defenders benefit more because they know their own systems — is weaker here than elsewhere. Defenders have to find and fix everything; attackers need one flaw. A tool that finds flaws at scale helps whichever side has the lower bar to clear.
This connects directly to what infrastructure security researchers have been reporting: commodity open-weight models are already the practical threat to industrial systems, ahead of frontier models, precisely because they are available.
What to do with this
Treat the benchmark as unconfirmed and the direction as confirmed.
Assume vulnerability-discovery capability is commoditizing — whether GLM-5.3 specifically hits 84.5% matters far less than open models arriving in that range at all. Run it against your own code before someone else does, because they will. Shorten your patch-cycle assumptions, since the window between a vulnerability being findable and being found keeps compressing. And watch the two-week safety review: what Z.ai does or does not restrict at public release is a useful read on how far open-weight labs are converging on frontier-lab safety practice.
For context, Z.ai is a Chinese challenger gaining real traction among Western developers, and results like this are how that traction gets built — match a restricted Western model, then ship the weights. The number may not survive independent testing. The strategy does not depend on it.
Source: China's Z.ai says new model nears Anthropic's Mythos 5 in cyber-defence tests — Reuters