Founder of Z.ai and Tsinghua Professor Jie Tang on scaling LLM

binyu2 pts0 comments

steve hsu on X: "Founder of https://t.co/J1fPFwBhJf and Tsinghua Professor Jie Tang:

"Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training." / X<br>Post

Log inSign up

Post

steve hsu on X: "Founder of https://t.co/J1fPFwBhJf and Tsinghua Professor Jie Tang:

"Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training."

steve hsu

@hsu_steve

Founder of Z.ai and Tsinghua Professor Jie Tang:

"Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time..."

Jie Tang quotes Einstein in his X profile:<br>“The value of a man should be seen in what he gives and not in what he is able to receive.”

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">jietang

@jietang

13h

Thoughts About Scaling Law

Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, Show more

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">12:18 PM · Aug 19, 202652.6KViews

42<br>489<br>202

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">steve hsu

@hsu_steve

6h

Z.ai built its own infrastructure for long horizon post-training.

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">steve hsu

@hsu_steve

Aug 16

GLM 5.3 shows power of automated long horizon post-training: frontier level agent and coding evals despite Z.ai: As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment. A useful Show more

1.9K

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">A War

@AWar1586398

1h

Yesss… the limit of parameter counts will top around 10T with transformer architecture.

96

Log in or sign up for X<br>See what’s happening and join the conversation<br>Continue with phoneContinue with AppleContinue with Google<br>or<br>Log in with username or email

Relevant people

steve hsu@hsu_steveFollow<br>Physicist, AI Founder, Manifold Podcast

Trending now

span empty before scaling post training

Related Articles