Z.ai releases GLM-5.3 open weights, claiming post-training gains in coding and cyber tasks
- GLM-5.3 is released as an open-weights model, and Z.ai says it uses the same base model as GLM-5.2, with all improvements coming from post-training.
- Z.ai reports Terminal Bench 3.0 at 28.3 for GLM-5.3 versus 4.6 for GLM-5.2, while GPT-5.6 Sol leads the table at 34.6 and Fable 5 with fallback scores 33.7.
- On cyber evaluations, GLM-5.3 scores 84.5 on CyberGym versus 77.2 for GLM-5.2, and Z.ai reports ExploitGym scores of 105 and 130 at 2-hour and 6-hour limits versus 29 and 39.
- The model supports local deployment through SGLang, vLLM, Transformers, KTransformers, Unsloth, TokenSpeed, and Ascend-focused inference frameworks.
- The reasoning_effort parameter accepts low, high, and max, defaults to max for unrecognized or omitted values, and the chat template defaults clear_thinking to false.
Hacker News 의견들
I've been using GLM models since December and have had a great experience without token drama or geopolitical restrictions. GLM-5.3 has me excited about eventually running this class of model at home.
I like that it skips Claude-style "load-bearing honesty" and just does the task. It is probably my favorite model to interact with, even if it is not always the most reliable.
GLM-5.3-Flash may be slightly more expensive than DeepSeek at list price, about $0.50 versus $0.48, though there is a temporary 50% discount. It has already gotten plenty of attention and provider support.
It has been very slow and inconsistent for me through z.ai. Neither GLM-5.3-Flash nor the new DeepSeek Flash felt flash-like, perhaps because they were overloaded.
My GLM-5.3-Flash tasks cost $0.30 or more where DeepSeek V4-Flash costs about $0.08. I used GLM High on OpenRouter with the presumed 50% discount, so I may be configuring something wrong, but it is slower too.
I used GLM 5.3 with pi and had a fairly good time. It is less restrictive on cyber work than US models and easier to run than Kimi, though I think Kimi is slightly more capable.
A used dual-Xeon or AMD server with 512GB RAM can run it locally, slowly. That can work for overnight analysis agents or one-shot modules, but the machine belongs in a garage or basement because it will be loud.
I built an EPYC system with 512GB of DDR4-3200 for about one-fifth the price of a 512GB Mac M5 Ultra. I want GLM as an architect and Qwen 27B or Next Flash as the implementer, probably at one-fifth the speed.
I would not buy expensive local-inference hardware based on tokens per dollar. But my pair of Sparks went from running GPT-OSS 120B with an AA score of 24 to GLM-5.3 Flash at Q4 with a score of 57, so the hardware has become far more useful over time.
Its reasoning disappoints me, even though it is good at solid grunt work for the price. It is bad at prose and conversation, and I cannot reliably tune away its habits, biases, and excessive back-and-forth.
I do not think building a state-of-the-art specialized model on top of it is feasible. General models keep beating specialized ones, and frontier releases may make your work obsolete before you finish it.
I am still interested in post-training an open model like GLM-5.3 with distillation and proprietary knowledge. The claim that specialist AIs can be 100x better makes me wonder whether there is training knowledge the public does not have.