Is AI Progress Real? Four Independent Metrics Show It

rbuccigrossi1 pts0 comments

Is AI Progress REAL? A SkepticCTO Analysis - SkepticCTO

SkepticCTO

SubscribeSign in

Is AI Progress REAL? A SkepticCTO Analysis

SkepticCTO<br>Jul 19, 2026

Share

Watch it here:

The following article includes the transcript and video sources.<br>Most attempts to graph AI improvement are a mess. This excellent graph by epoch.ai clearly shows the problem:

Credit: Epoch.ai (CC-BY)<br>Almost as quickly as benchmarks are defined, they are saturated. Are large language models truly getting better? Or are they learning to study to the tests?<br>So I took a journey to find measures of AI improvements that have lasted over the past couple of years. I found 4. Take a look at this:

Flat for years, then sharply rising. Here’s the thing: These are 4 different types of measures, yet they bend at the same point… the last quarter of 2025.<br>So what happened at the end of 2025?<br>November 17: xAI ships Grok 4.1 and tops the leaderboards.

The next day on November 18: Google ships Gemini 3, the first model to break 1,500 on the LMArena leaderboard.

Six days after that November 24: Anthropic releases Claude Opus 4.5.

Two weeks after that on December 11: OpenAI answers with GPT-5.2.

Four flagship models from four different labs. All within twenty-five days.<br>If you felt something shift in AI capability last winter, the graph agrees with you. And since then the capabilities seem to improve month after month. But does this graph represent real improvement? Will this trend continue? That’s what we’re going to explore in today’s news episode.<br>The Multiproxy Test

Here’s how you tell a real signal from a measurement artifact: check whether independent methods, run by independent people, tell the same story.<br>Any single benchmark curve is easy to dismiss, and for good reason. There’s contamination, where test data leaks into training data. There’s saturation, where a benchmark was too easy to begin with and everyone eventually caps out. There’s reward hacking, where a shortcut in the test design inflates scores without the model actually getting more capable. These are real problems, and they have the potential to cause us to question real phenomena.<br>Climate science ran into this exact problem in the late 1990s. Michael Mann, Raymond Bradley, and Malcolm Hughes published a northern-hemisphere temperature reconstruction in 1998 that became known, informally, as the hockey stick. Critics called it cherry-picked, a measurement artifact dressed up as a trend. What settled the argument wasn’t a better defense of the original chart. It was independent corroboration. Tree rings, ice cores, borehole temperatures, coral records, each with a different error mode, all pointing the same direction. When methods that can’t share the same mistake agree with each other, that strengthens the case for the underlying story.<br>With AI we have two extremes. One is unbridled excitement: treating any upward-sloping chart as proof of something profound. The other is reflexive dismissal: treating “benchmarks can be gamed” as a permanent excuse not to look at the data. Contamination and saturation are legitimate concerns. But they can’t be used to dismiss every chart without evidence.<br>So does the AI hockey stick hold up under the multiproxy test? Four independent metrics, measured by four different organizations and four different methods, all show the same bend at the same time.<br>Four Independent Measures

The first measure, METR’s time horizon, has the longest continuous track record. METR measures the longest software engineering task a frontier model can complete successfully half the time.

In March 2024, Claude 3 Opus held the record at four minutes. By February 2025, Claude 3.7 Sonnet had reached one hour. By November 2025, Claude Opus 4.5 was approaching five hours. By February 2026, Claude Opus 4.6 had crossed twelve.<br>Four minutes to twelve hours, in under two years.<br>The second measure, TrackingAI’s offline cognitive test is worth attention because it’s kept off the public internet, which limits training contamination. TrackingAI publishes IQ scores (which has a bell shaped curve). So I converted it into something more intuitive: one out of X, where X number of humans you’d need to test before finding one who beats the score.

In March 2024, more than nine out ten people outperformed the best AI system on this test (or 1 out of 1.1). By June 2026, only one out forty-four would.<br>The third measure, Humanity’s Last Exam is a tough exam with Ph.D. level questions across many different disciplines. A 2025 paper posted to OpenReview estimated that the compute needed to hit specific HLE scores is exponential. So to get from the score of 2.7% in May of 2025 to 53% in June of 2026 represents 200 times as much compute effort.

The fourth measure, ARC-AGI-2 tests fluid, abstract reasoning instead of recalled knowledge, and its leaderboard tells a pretty strange story. From May 2024 through August 2025, GPT-4o and GPT-5 sat at 9% and 9.4%, essentially flat for fifteen...

four real test independent different skepticcto

Related Articles