Z.ai says GLM-5.3 more than doubles GLM-5.2 on exploit benchmarks using post-training alone
- Z.ai says GLM-5.3 uses the same base model as GLM-5.2 and gains entirely from scaled post-training, reporting a 50% improvement on its internal Z.ai Code Bench.
- On CyberGym, GLM-5.3 scores 84.5 versus GLM-5.2's 77.2; Z.ai reports ExploitGym scores of 105 and 130 at two and six hours, versus 29 and 39 for GLM-5.2, plus 54.4 versus 24.4 on ExploitBench.
- GLM-5.3 reaches 28.3 on Terminal-Bench 3.0, up from 4.6 for GLM-5.2, and scores 66.9 versus 46.2 on DeepSWE v1.1; its published comparison table still places some closed models ahead on several coding and cyber tests.
- Z.ai builds training environments with research agents that turn work patterns into executable long-horizon tasks, then uses a judge agent and synthesized verifiers to check solvability and create binary RL rewards.
- Z.ai says it will publish GLM-5.3 weights two weeks after launch, after safety evaluation and hardening.
Hacker News opinions
It is hard to tell which of this flood of models to use unless you test them on complex real work yourself. Price is about the only simple filter.
I am deciding whether to buy a Z.ai coding plan or wait for OpenRouter. GLM-5.2 was solid in my security-auditing tests, but a little behind Opus 4.8, so I mostly use Kimi K3 because US vendors restrict security work.
The numbers look incredible, but I will wait to see real-world performance. GLM release timing always catches me by surprise.
OpenAI and Anthropic should give people access to their cyber models. Attackers can use open and closed models, while maintainers who depend on those vendors may not get approval to use equivalent tools.
I have to switch to Kimi or GLM even for basic issue triage in my own projects. The current guardrails are ridiculous.
GLM-5.3 still looks a little behind Sol and Fable, but only narrowly. I do not see a strong economic reason to leave OpenAI yet, though post-training on the GLM-5.2 base is getting very close.
I use DwarfStar for GLM-5.2 and DeepSeek. It is useful for actual work, not only experimentation.
For security work, the closed-model comparison barely matters to me if Fable and Mythos will refuse the request. Normal users cannot access those capabilities anyway.
Open models are much more useful for some work because jailbreaking is trivial if you know how. In reverse engineering, I have seen GLM-5.2 talk itself into treating a task as a crackme or CTF without any prompt from me.
Even with a small capability gap, a free model changes the economics. That alone can beat Sol and Fable for many users.
I would expect at least two DGX Sparks for a 2-bit quant, and quantization damages these models badly. Four Sparks may run GLM-5.2 at 4-bit, but I would rather run DeepSeek-V4 Flash at full precision.
Chinese labs are producing open-weight models that US providers can host cheaply and monetize. I do not see how US labs justify trillion-dollar valuations if model capability commoditizes this quickly.
I wish the weights came under a true FOSS license. Kimi and Qwen moving to restricted-use licenses is a retreat from the earlier open-source Chinese LLM culture.
Freely available weights for indie developers, small companies, and research may be the financially sustainable compromise. Companies operating thousands of GPUs and earning millions should pay.
The opening claim that scaling post-training was all they did is striking. If the base is unchanged, I want to know whether GLM-5.3 has the same parameter count as GLM-5.2 and what "scaling post-training" means in practice.
It seems like both parameter count and post-training matter. GLM-5.3 competing with models reportedly three to four times larger suggests parameter count is no longer a direct proxy for performance.