AI Has Passed Every Exam. It Has Never Had an Idea.
Manas Bihani
SubscribeSign in
AI Has Passed Every Exam. It Has Never Had an Idea.
Manas Bihani<br>Aug 01, 2026
Share
The machine has passed every exam we can build and discovered nothing. Both are true. What dissolves the contradiction is the one thing no test has ever measured in machines, or in us.<br>Here are two facts about AI in 2026 that you have almost certainly seen, but never had to hold in the same hand.<br>Thanks for reading! Subscribe for free to receive new posts and support my work.
Subscribe
The first: the machine passes everything. Humanity’s Last Exam, a wall of graduate-level questions built specifically to be too hard, it clears. Gold at the International Math Olympiad. The bar exam, the medical boards, competition-grade code. We keep building harder tests, grandly named, designed by PhDs to find the ceiling, and the machine keeps stepping over them. On any exam a human can be given, it is now roughly the best test-taker alive. Frontier models gained 30 percentage points in a single year on Humanity’s Last Exam alone.<br>The second is harder to state cleanly than it was even a year ago. AI systems now genuinely contribute to scientific discoveries. AlphaFold, which won the 2024 Nobel Prize in Chemistry, GNoME, which predicted 2.2 million crystal structures of which outside labs have physically synthesized 736, FunSearch and others have shown that much. But notice the shape of those discoveries. In every case, a human decided the question was worth asking. The machine answered it, often spectacularly. I could not find an example where it independently decided which question mattered. No AI has posed a question nobody thought to ask, found an anomaly nobody was looking at, or made the kind of leap that turns a field inside out. It answers our questions superbly. It has never, on its own, decided which question was worth asking.<br>Both true. Sit in the discomfort for a second, because the discomfort is the whole essay. The best test-taker in history, and it has never had an idea. If intelligence were one thing, that shouldn’t be possible.<br>So what mechanism would reliably produce this outcome, over and over, regardless of how good the underlying system gets?<br>The easy dissolves don’t work
Everyone reaches for a comfortable way out of that contradiction, and both exits are blocked.<br>The first exit: it’s not really thinking, it’s just autocomplete. That feels good and explains nothing, because “just autocomplete” doesn’t clear Humanity’s Last Exam. Whatever it’s doing, dismissing it doesn’t survive the scoreboard.<br>The second exit: give it time and scale, discovery is coming. Maybe. But that’s a promise, not an explanation, and it dodges the actual question which is why the gap has this particular shape. Why superhuman on every answer and silent on every question? Scale doesn’t explain a structural asymmetry. It just promises to wash it away later.<br>Around halfway through writing this essay, Tom Zahavy at Google DeepMind published a position paper called LLMs Can’t Jump. For about ten minutes I thought I’d been beaten to the idea. Then I realized we were asking different questions. His argument runs through Peirce: models handle induction and deduction, but not the abductive leap that invents an explanation the data never contained. He has since clarified that this is a personal position rather than DeepMind’s, and that scaling might prove him wrong.<br>Zahavy asks whether the machine can make the leap. I found myself asking something slightly stranger: suppose it did. How would our benchmarks know?<br>So there’s a hidden variable, the way there always is when two careful observations point in opposite directions. And the variable isn’t in the machine. It’s in the tests.<br>A benchmark is a question someone already chose
Here is the thing that dissolves the paradox, and once you see it you can’t unsee it.<br>Every benchmark, every exam, every one of those grandly-named tests, has the same structure: someone writes the question, and the machine finds the answer. The question is given. It arrives pre-selected, pre-formatted, flagged as important, with a known answer sitting in a locked drawer so the thing can be graded.<br>And that means every test we have ever built measures exactly one half of intelligence, the finding of answers and structurally cannot measure the other half: the choosing of the question. A test can’t measure question-choosing, because a test is a chosen question. The container can’t weigh the thing that decides what goes in the container.<br>That sounds like wordplay until you look at where the great leaps actually came from, and notice that the choosing was always the hard part.<br>Einstein didn’t answer the question. He found it.
In 1887, two physicists named Michelson and Morley ran an experiment expecting to measure how the Earth’s motion changed the speed of light. They got nothing. Light moved at the same speed no matter which way you chased...