Announcement post published on August 14, 2026 on the Z.ai blog (formerly Zhipu AI), unsigned, for the launch of GLM-5.3.

The methodological thesis. « Scaling post-training is all we did for GLM-5.3. » Same base model as GLM-5.2: all the gain comes from post-training, built on the stack from the previous cycle — IndexShare (long context), SAO (long-horizon RL) and slime (asynchronous training, Megatron + SGLang). The bottleneck has shifted from the model to the environment: Z.ai describes pipelines that synthesize environments and reward signal — a judge agent verifies solvability, verifiers are synthesized without access to the reference solution, and are admitted only after a triptych of oracle / no-op / unsolved-state controls. The work remains « human-in-the-loop ». End-to-end RL throughput improved by more than 2.3×.

As part of post-training, we introduced vulnerability discovery data and environments into the training mix. We expected this to make the model better at finding and reasoning about vulnerabilities

**Z.ai** , z.ai

The coding results. Terminal-Bench 3.0 goes from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, Agents' Last Exam from 23.8 to 28.5. On Z.ai Code Bench, an in-house, private benchmark, +50% over GLM-5.2, with a simultaneous gain in token efficiency: 34.5% at ~75K output tokens at Max effort (versus 23.4% at 96K for GLM-5.2), and 31.4% at ~50K at High effort — ahead of Claude Opus 4.8 (29.5% at 120K). Claude Fable 5 remains ahead at 39.5%. The claim « most capable open-weights model for coding » does not follow from the table: against Kimi K3, the score is 3–3 with one tie.

The cyber capability. Presented as « emergent », it was deliberately trained — the post writes « we expected this to make the model better ». What came as a surprise was the speed, and the shift from isolated flaws to the complete exploitation chain. CyberGym 84.5% (best in the table), ExploitBench 54.4% (×2.2), ExploitGym 105/130 tasks (×3.6 over GLM-5.2, throughput-normalized budgets). Key sentence: « Capability is growing fastest exactly where we are furthest behind. »

The heaviest number. Working with Chinese security teams, the model identified 2,436 vulnerabilities in 269 open source projects — kernels, OSes, browser engines, network protocols — the oldest introduced in 1981, average lifetime 26.6 years. The Security Disclosure Ledger shows 53 disclosed and 2,383 under embargo: 2.2% published.

Governance. Weights announced « in two weeks, once safety evaluation and hardening are complete »a date, not a criterion: no definition of hardening, no condition for non-release, no third-party evaluator.

Miscellaneous. thinking.type: "disabled" is no longer supported (migration required); GLM Coding Plan quotas in points, 50% outside 14:00–18:00 UTC+8; nearly all evaluations are conducted in Claude Code 2.1.207.