Convergence in LLM Quality and Slowdown in LLM Improvement

Anon841 pts0 comments

Convergence in LLM Quality & Slowdown in LLM Improvement: CHART OF THE DAY

SubscribeSign in

SubTuringBradBot<br>Convergence in LLM Quality & Slowdown in LLM Improvement: CHART OF THE DAY<br>From two Epoch-Capability points a month back in 2023 to one every six months today, with the spread of assessed model capabilities across frontier labs shrinking by half...

Brad DeLong<br>Jul 21, 2026

77

13

Share

But are these http://epoch.ai> assessments real numbers that mean anything? And how could we tell? The apparent slowdown in LLM improvement is exactly what you would expect if the LLMs are at base just emulating internet s***posters. But if the compression = true understanding crowd is right, the scale may well be measuring the wrong thing…<br>Whether we are on the road to AGI hinges on an unsettled question: What are LLMs?<br>If they are sophisticated mimics of human conversation, plateauing capability scores make sense and the AGI story is hype.<br>If compression-into-weights actually recovers, by seeking minimum Kolmogorov complexity, the true generating structure of deep thought that humans carry on before they then produce their jittery and very human text, then maybe it is time to start planning to welcome our potential AGI overlords.<br>Paul Kedrosky reprints a graph he has had on his mind for some time:

Paul Kedrosky : Kimi, Model Convergence, & the Post-Training Era https://paulkedrosky.com/kimi-model-convergence-and-the-post-training-era/>: ‘Moonshot’s Kimi K3 model has people over-excited, as if some trend has been broken, but they’re wrong…. We exited the pre-training era and became more reliant on post-training, like RLHF. In the post-training era, successive releases deliver smaller capability gains, fewer durable outliers, and less defensible technical differentiation…. Te value of each incremental model release (ignoring harnesses) is falling, even if production costs aren’t…. Model prices compress…. Inference becomes increasingly commoditized…. Frontier development becomes harder to monetize…. Value shifts away from the base model…

The big joker in his analysis of course is this: What exactly is on the vertical axis of these scores from http://epoch.ai>? Why should we care? What difference does it really make?<br>If you believe, as I do, that LLMs are going around Robin Hood’s barn, yeah, because they’re “emulations of the typical internet shit poster”, then frantically corrected to sanity by RLHF and such, this is as expected. You are trying to faithfully emulate what the person whose ghost you are pantomiming said, so it is really hard to get smarter than them. After a point, the better classification of what human conversations are “close” to the one the LLM is having is not worth much at all.<br>That leaves the use of MAMLMs for very big-data, very high-dimension, very flexible-function classification. True deep magic appears to have emerged with respect to programming here—perhaps. And there is no doubt that other superuse cases will emerge at a rate and with an impact I cannot forecast.<br>There are, however, people who say that the frontier model-builders are closing in on true “AGI”, true “Artificial General Intelligence”, something that, surveying across all of the subdomains of cognition and averaging, is human-level albeit not human-like.

When I ask them how it can possibly do this, they are either (a) silent, or (b) they say compression. They say that the “fuzzy .jpeg of the web” line of Ted Chiang’s gets it exactly backwards. What the LLM does as it compresses its training data into the weights of its virtual emulation of a neural network is to, in some way, generalize. What it loses is not fuzziness, but the jitteryness of individual human error. Thus it gains the wisdom of crowds à la Sean Trott https://seantrott.substack.com/p/gpt-4-sometimes-captures-the-wisdom>: it says not what the human did in the closest conversation, but what each of the humans would have said in all of the close conversations if each had had knowledge of what all the other humans in similar conversations were saying and thinking.<br>What that does is produce structural recovery: true knowledge of what is going on, in a context where it is the humans who each have only a fuzzy .jpeg view of the surface appearance of things, or at least of relationships between words.<br>If they are right, then this https://epoch.ai/eci?subset-view=graph&subset-tab=Software+engineering&view=graph&tab=release-date> Capability Index is fundamentally wrong, and misleading us.<br>Um—maybe?<br>How could I figure this out? Damned if I know. Some thoughts below:<br>As best as I can figure it out, the pro-compression-induces-true-knowledge line is pushed by people like Ilya Sutskever https://www.youtube.com/watch?v=AKMuA_TVz3A> https://www.ted.com/talks/ilya_sutskever_the_exciting_perilous_journey_toward_agi> and Jack Rae https://www.youtube.com/watch?v=dO4TPJkeaaU>: Rae, especially, sees LLM training as computing the probability distribution that lets you...

model true human training https convergence

Related Articles