Z.ai has introduced GLM-5.3, a model the company positions as “Built to Code. Ready for Cyber Defense.” The release focuses on coding, agentic tasks and cybersecurity, with Z.ai claiming that the model improves substantially on GLM-5.2 while producing stronger results with fewer output tokens.
Z.ai launches GLM-5.3 with a focus on coding and cybersecurity
GLM-5.3 was post-trained on a 743-billion-parameter base model. Z.ai describes it as having “top-tier” coding and agentic capabilities, as well as a major cybersecurity advance among open models. Those claims are supported primarily by the company’s own benchmark results and should be treated accordingly until independent testing is available.
The model is currently accessible through the GLM Coding Plan and ZCode. Z.ai plans to release API access and open weights in stages after additional safety evaluations. The company also mentions an initial group of partners offering GLM-5.3-powered services through its official service, with the partners operating under Z.ai’s safeguards and usage policies.
Benchmark results show gains over GLM-5.2
Z.ai’s Code Bench v1.0 evaluation was conducted with Claude Code 2.1.207. GLM-5.3 scored higher than GLM-5.2 across most of the listed coding, cybersecurity and agentic benchmarks.
Selected results include:
| Benchmark | GLM-5.3 | GLM-5.2 |
|---|---|---|
| Terminal Bench 3.0 | 28.3 | — |
| DeepSWE v1.1 | 66.9 | — |
| NL2Repo | 58.0 | — |
| ProgramBench | 19.0 | — |
| FrontierSWE | 78.1 | — |
| SWE-Marathon v1.1 | 42.5 | — |
| PostTrainBench | 39.8 | — |
| CyberGym | 84.5 | 77.2 |
| AutomationBench v1.0.6 | 48.2 | 26.2 |
The full benchmark set also includes a score of 88.2 on Terminal Bench 2.1, 73.0 on Toolathon Verified, 28.5 on Agents’ Last Exam, 62.5 on HLE with Tools and 1,769 on GDPval-AA v2.
GLM-5.3’s CyberGym score of 84.5 was higher than the listed results for GLM-5.2, Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, Opus 4.8, Fable 5 and GPT-5.6 Sol. On AutomationBench, its 48.2 score also exceeded the listed scores for those comparison models.
The model did not lead every comparison. Its Toolathon Verified score of 73.0 trailed Kimi K3 at 76.5 and DeepSeek-V4 Pro-0813 at 74.1. Its Terminal Bench 2.1 result of 88.2 was close to Kimi K3’s 88.3 and below GPT-5.6 Sol’s 88.8. On DeepSWE, GLM-5.3 scored 66.9, compared with 67.5 for Kimi K3, 69.7 for Fable 5 and 72.7 for GPT-5.6 Sol.
Partner testing points to lower token use
Command Code, one of the partners working with GLM-5.3, reports that the model escaped all 15 deliberately induced loops in an internal evaluation. The company described the test as small, so the result is an early observation rather than a broad measurement of reliability.
Command Code also reports approximately 50–60% token savings on an average session when GLM-5.3 is paired with its tool-defer harness. That finding is specific to Command Code’s tooling and workload, and may not transfer directly to other coding environments.
Z.ai’s launch materials emphasize a chart comparing GLM-5.3 with GLM-5.2, Claude Fable 5 and Claude Opus 4.8 across low, high and maximum effort levels. The comparison measures both accuracy and average output tokens per task, supporting the company’s claim that the newer model can improve results while generating fewer tokens in the tested scenarios.
Broader evaluation will depend on API availability, open-weight releases and independent tests across different hardware and coding setups. For now, GLM-5.3 is available through Z.ai’s own coding products, while wider access remains contingent on the company’s safety review process.
Source: Z.ai on X