Does Forecasting Have Room At The Top? - by Scott Alexander
Astral Codex Ten
SubscribeSign in
Does Forecasting Have Room At The Top?<br>...
Scott Alexander<br>Aug 03, 2026
108
81
Share
Superforecasting is the art/science/sport of predicting the future - for example, who will win elections, which countries will fight wars, when key technologies will be discovered. Over the past few years, it went from an obscure academic subfield to a multibillion dollar industry in the form of prediction markets. More recently, AI superforecasters have come close to the accuracy of top humans, and their performance is rising rapidly. In a year or two, we’ll see one of the following patterns:
Either humans have already come close to some fundamental limit on the predictability of world events - in which case AIs will plateau at or slightly above the human level - or the trend line will continue until AIs are far beyond top humans. By analogy to superintelligence, the natural term for the second situation would be “superforecasting”; since that’s already in use, we can cringely call it “ultraforecasting”.<br>Daniel Reeves1 makes the case for scenario A here. He describes a study he coauthored in 2010, which found that, on a variety of questions related to sports games and movie box office receipts2, prediction markets only outperformed simple boring statistical models by 3-6%. Maybe those statistical models are close to the best that it’s possible to do; the rest is what the mathematicians call aleatoric uncertainty - irreducible complexity downstream of chaotic systems that entirely resist modeling.<br>He could be right. This post isn’t meant to be a decisive refutation, but rather a description of why I’m still about 70-30 expecting Scenario B.<br>Slightly Contra Goel, Reeves, et al
Reeves’ study claims that the prediction markets of 2010 only beat dumb statistical models by 3% (for sports) to 6% (for movies).<br>But these percentages aren’t real win-loss percentages; they’re variation in a quantity called root mean-squared error. One way to get a feel for this quantity is that the dumb statistical model for sports (home team advantage + win-loss record) beat an even dumber statistical model (home team advantage only) by 0.8 percentage points, and the prediction market beat the first (better) statistical model by another 0.4 pp. So the effect of going from a statistical model to a prediction market is half as large as the effect of knowing which two teams were playing and how good they are! On this metric, the prediction markets are a vast improvement over previous state-of-the-art.
The benefit of prediction markets over statistical models is half as great as the benefit of including a team’s previous win-loss record in a statistical model which previously didn’t have that! (source: Reeves et al Figure 2, Claude Fable)<br>Why can we frame this same result as either very small or very large?<br>Sports are optimized against prediction3. If there were a fully predictable sport (eg heightball, where all athletes line up in a row and the tallest one wins), nobody would watch it. Instead, we go to absurd lengths to keep the outcome uncertain. Salary caps, draft systems, etc try to ensure that all teams have exactly equal talent. Commercial incentives and ceiling effects ensure that they have exactly equal training. Then an exactly-equal number of these exactly-equally-talented-and-trained people are placed in exactly-identical positions on a perfectly-symmetrical field and told to hit/kick/throw a ball which is placed exactly equidistant between both of them. It’s funny for me to describe it this way, because obviously this is what we want (to “keep things fair”), but it’s all designed for prediction-resistance. Given the difficulty of the domain, even very large relative advances in prediction look small in absolute terms.<br>Maybe we should look at Reeves’ other example, movie box office receipts. Here the markets did slightly better, getting a 6% improvement. But isn’t this still low?<br>Box office receipts differ by orders of magnitude (some movies make $100,000, others make $100 million), so the paper puts this on a log scale. 6% improvement on a log scale is already starting to sound pretty good.<br>And again, it all ends up coming down to what we compare it to. Here the super-dumb model is that all movies make $8.1 million, the takings of the exact average movie. The slightly-less-dumb model then adds the number of screens that the movie is showing on and the amount of Google search traffic for the movie! For example, a random indie film might be showing on three screens in the entire country, and a Disney blockbuster might be showing on ten thousand. The challenge the paper gives prediction markets is to significantly improve on knowing whether a film is an indie film or a Disney blockbuster, plus knowing how many people are interested in seeing it, and it has to do this on a log scale! No wonder the relative improvement number comes out...