GLM-5.3: How Chinese labs keep stride with the frontier
SubscribeSign in
GLM-5.3: How Chinese labs keep stride with the frontier<br>Hint: It’s really not a distillation story.<br>Nathan Lambert<br>Aug 14, 2026
65
Share
Housekeeping: I’m traveling so cannot make a voiceover for this post. EDIT — I added a bullet point 5 on the Chinese data industry after sending the email out.<br>Today, Z.ai announced their GLM-5.3 model, currently only available in the coding plan, coming soon to their API and in two weeks’ time to Hugging Face (open weights). This model looks exceptional, with a somewhat astounding increase in scores. On many benchmarks the model has surpassed Moonshot AI’s Kimi K3 and on some it’s surpassed Claude Fable 5 or GPT-5.6-Sol.
Here’s a more complete comparison:
This puts the model more or less at the frontier of agentic coding benchmarks, with only ~750B parameters – a third of Kimi K3! The Z.ai blog post is rather straightforward, and starts with a bold sentence:<br>Scaling post-training is all we did for GLM-5.3.
GLM-5.3 is the same base model as GLM-5.2 with substantially extended post-training. To risk a broad oversimplification, Z.ai seems to have a strength in post-training when compared to Kimi, which is more of a pretraining masterpiece. Following this release there have been a lot of discussions wondering how China can keep up so well? How can such a small model be matching the leading public American models? Are these results real?
Subscribe
The simplest explanation is that Z.ai is very good at what they do – it’s worth recalling that they’ve been working on this line of models longer than almost anyone in the industry. Here’s a brief history of the GLM models.<br>Zhipu AI Founded – 2019
GLM (General Language Model) — March 2021 — released by THUDM , Tsinghua University’s Data Mining / Knowledge Engineering group. Weights
GLM-130B — August 2022 — Scaled version. Technical report for GLM-130B through GLM-4 — Weights
ChatGLM — March 14, 2023 — first chat version. Weights
ChatGLM2 — June 25, 2023 — Weights
ChatGLM3 — October 27, 2023 — Weights
GLM-4 — January 16, 2024 — rebranded as just GLM; open-weight GLM-4-9B followed in June. Weights
GLM-5 — February 11, 2026 — latest major generation. Weights
GLM 5.2, released on June 22 of this year, was a big deal – weeks after the release, I regularly heard from AI researchers I know who still used the model due to its speed (some deploy the model on internal clusters for faster speeds than public offerings) and simplicity (as a model with no rollbacks, etc., when working on frontier AI systems). GLM-5.2 altogether stood up to the hype.<br>I’ve been going through some of the same denial myself, thinking “how do they keep doing this? Surely the models aren’t as good as they look.” There’s something a bit off-putting with how the American companies have such a commanding resource lead, but can’t seem to pull away in capabilities. The common answer is distillation, which I’ve written at length about, but I deem not to be the major factor. On that note, there was a recent paper that showed simple methods for extracting the reasoning traces from frontier models – this is the sort of thing that Chinese labs could definitely use at scale. I’m confused why the labs in the U.S. haven’t patched this behavior faster; instead they’re running to the government asking for policy help. It doesn’t add up for me.<br>Z.ai’s blog is direct and matches with an RL-dominated training regime. They say they used “more environments, more diverse tasks, and more compute spent training on them.” One does not simply “distill” RL environments, infrastructure to run them at scale, or algorithms to mix them together effectively.<br>Interconnects AI is a reader-supported publication. Consider becoming a subscriber.
Subscribe
So, how do the Chinese labs do it if not distillation? Are they benchmaxxing? An accepted definition of benchmaxxing is focusing the model on the test sets, such that the real-world performance meaningfully differs from the on-paper scores. The determining factors are much more big picture than technical (yes, the technical details definitely matter, but are harder to differentiate from lab to lab):<br>The time to release for Z.ai is likely days, not months as with OpenAI or Anthropic. It is very, very likely that OpenAI and Anthropic have far better internal models than Z.ai and Moonshot AI. Still, these American companies tend to take months to release their models to the public, which massively flatters the Chinese labs in adoption decisions at the frontier. To put it simply – the Chinese labs use all the time that American labs do pre-release testing to keep hillclimbing on benchmarks (SpaceXAI is likely far closer to the Chinese labs here). With the pace of progress being so fast, this is likely the largest determining factor of why Chinese labs stay at the frontier. This, so far, has been economically acceptable for the American labs, as they’ve still had...