Zhipu AI, the Beijing-based lab behind the GLM model family, released GLM-5.3 on August 14, positioning it as a coding-focused update to June’s GLM-5.2. The bigger story, according to the company’s own announcement, is what the model turned out to be unexpectedly good at: finding and exploiting software vulnerabilities.
Coding gains from post-training alone
Zhipu says every capability jump in GLM-5.3 comes from extended reinforcement-learning post-training on the same base network as GLM-5.2 — the architecture itself didn’t change. The company reports a 50% improvement on its internal coding benchmark, with GLM-5.3 climbing from 4.6 to 28.3 points on Terminal-Bench 3.0 and from 46.2 to 66.9 on DeepSWE v1.1, putting it first among open-weight models on both and, Zhipu claims, within reach of Anthropic’s flagship Claude Fable 5 on coding tasks.
An unplanned jump in cyber capability
The same post-training scaling produced a sharper rise in offensive-security skill than Zhipu expected. GLM-5.3’s score on the CyberGym benchmark rose from 77.2% to 84.5%, edging past Anthropic’s Mythos 5 (83.8%) and OpenAI’s GPT-5.6 Sol (83.6%), though it still trails both on the harder ExploitBench test despite nearly tripling its own score there, from 24.4% to 54.4%.
In real-world testing run with outside security researchers, GLM-5.3 surfaced 2,436 vulnerabilities across 269 open-source projects, Zhipu said, including 1,097 rated medium-to-high severity. The oldest flaw traced back to 1981, and the average bug had gone undetected for 26.6 years. Fifty-three of the findings have already been disclosed publicly; the rest remain under embargo while Zhipu coordinates fixes with maintainers through a new tracking system it calls the Security Disclosure Ledger.
Weights held back for “hardening”
GLM-5.3 is live now through Zhipu’s GLM Coding Plan and API, and works with third-party coding agents such as Claude Code and OpenCode. But unlike most of the company’s past releases, the open-weight version won’t ship immediately — Zhipu says it needs roughly two weeks for “safety evaluation and hardening” first. “Cyber capability developed faster than we expected,” the company wrote, adding that capability is “growing fastest exactly where we are furthest behind” relative to Western labs.
The delay puts Zhipu alongside OpenAI and Anthropic, which have both recently tightened review for newer models after they crossed internal cybersecurity thresholds — though GLM-5.3 stands out for being an open-weight model whose weights, once released, anyone will be able to download and run without Zhipu’s own safeguards attached.