SemiAnalysis: Are Open Models Catching Up?

j-bu2 pts0 comments

Are Open Models Catching Up?

SubscribeSign in

Are Open Models Catching Up?<br>Comparing open vs. closed models across the eras of frontier models, Is the gap narrowing?<br>Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel<br>Aug 21, 2026<br>∙ Paid

41

Share

The past two months have been a breakout period for open source AI. Yes, there was the “DeepSeek moment” back in January 2025, but no one actually used R1 to do any economically valuable work. In contrast, models like GLM 5.3 and Kimi K3 are genuinely capable of many of the same coding and agentic tasks that rocketed Anthropic to $65B+ ARR . Unlike others who inflated ARR, our figures were much closer to reality.

Source: SemiAnalysis<br>It is an exciting time to be a token consumer. Competition is heating up, usage resets are being doled out, and the battle for your tokens now extends beyond the OpenAI-Anthropic duopoly. Fireworks alone is processing over 40T tokens per day—2x the OpenAI API’s volume at the end of March.<br>However, major FUD has also emerged as a result of open model success : if open models stay capable enough relative to the closed frontier at a fraction of the cost, won't the model layer become commoditized? This outcome would obviously be disastrous for frontier lab margins. For full details on Anthropic and OpenAI’s financials, see our Tokenomics Model.<br>To project how the open vs closed capability gap will progress in the future, we first need to measure the past. Naively, you might pick a single set of benchmarks to measure all historical models, but this is a mistake. Every benchmark is a product of a particular era. When someone creates a new benchmark, their goal is to discern differences in model capabilities at the time. If they’re successful, the model makers will climb said benchmark until it becomes saturated. Once that happens, everyone stops caring about the benchmark, and the cycle repeats.<br>There have been three eras thus far in the history of LLMs: early scaling, reasoning, and agentic . Each era represented a step-function increase in model utility, and rather than trying to plot a single continuous trend, we believe it’s better to evaluate the models and benchmarks from each era individually.<br>When viewed this way, it becomes clear that the open vs. closed gap moves in cycles . At the start of each era, a frontier lab completes some promising research, trains an impressive model, deploys it at scale to their users, and jumps ahead. Then, other labs identify the key advances, reverse-engineer what the frontier lab is doing, replicate them in their own models, and close the gap. Nothing stays secret forever—especially when you factor in distillation. It’s just a question of how long it takes.<br>To answer this question, we took all the relevant models from each era and ran a curated set of benchmarks to get a composite capability score. The result is a clear trend: with each generation, open-source models take half as long to catch up to the first closed-source model of the era.

Source: SemiAnalysis<br>Of course, benchmarks don’t tell the full story, and we’ll highlight all the relevant caveats below. Finally, we’ll extend this analysis into the future, and explain why it’s less bearish frontier models than you might initially think.<br>How we measured

Here’s an overview of the models and benchmarks we selected for each era:

Source: SemiAnalysis<br>Picking a single SOTA closed and open model at a particular time is subjective, but our selections reflect the general consensus among AI experts. In cases where there’s debate—e.g. Fable 5 vs GPT 5.6 today—we were conservative and tested both.<br>For benchmarks, we relied on a combination of personal taste and popularity. Humanities Last Exam (HLE), for example, is known to have lots of issues, but was also truly one of the defining benchmarks of the reasoning era with no close substitutes. SWE-bench Pro, on the other hand, is similarly popular and problematic, but also closely approximated by DeepSWE.<br>Most of the benchmark scores here we ran ourselves using Prime Intellect's evaluation stack, specifically their environments hub and the evals harness included in Prime-RL. The rest come from runs by our friends at Artificial Analysis and Datacurve's DeepSWE leaderboard. Open models were served the way they would have been at release: vLLM versions, hardware that was in use at the time, and sampling settings from the model card. For closed models, we ran against their pinned API versions. Where our numbers share a chart with third-party values, we matched their rulesets.<br>We’d like to give a huge thank you to Florian Brand (@xeophon) from Prime Intellect for helping us pick benchmarks/models, implement evals, and check for correctness.<br>Era 1 | Early scaling (2022-2024)

It’s June 2023. The world is reckoning with ChatGPT, and Mark Zuckerberg just agreed to fight Elon Musk at the Colosseum. But while Zuck is training jiu-jitsu and doing Murphs, his company is doing some training of their own. FAIR...

models open model benchmarks closed frontier

Related Articles