Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazine
About Quanta
Search
Search for:
Search<br>Search
Newsletter
Get the latest news delivered to your inbox.
Subscribe
Recent newsletters
Follow Quanta
Youtube
RSS
An editorially independent publication supported by the Simons Foundation.
Type search term(s) and press enter
What are you looking for?
Search
Home
Is AI Reasoning Right for the Wrong Reasons?
Comment
Save Article
Read Later
Share
Copied!
Copy link
Ycombinator
Comment
Comments
Save Article<br>Read Later
Read Later
Qualia
Is AI Reasoning Right for the Wrong Reasons?
By
John Pavlus
July 31, 2026
The idea that artificial intelligence can “reason” is more intuitive than ever. But intuitions can be wrong, and the science is far from settled.
Comment
Save Article
Read Later
Celsius Pictor for Quanta Magazine
Introduction
By John Pavlus
Contributing Writer
July 31, 2026
View PDF/Print Mode
artificial intelligence
computer science
features
large language models
machine learning
Qualia
All topics
I’ll just say it: What the hell is going on with AI “reasoning”?
Sorry for the air quotes. That punctuational side-eye was more common in 2024, when the specially trained cousins of LLMs now known as “large reasoning models,” or LRMs, were still new. Nowadays it may seem downright churlish, though, given that a “general-purpose reasoning model” from OpenAI solved a famous open mathematical research problem in one shot in May 2026. Still, I’m not sure how else to acknowledge my intellectual whiplash over the scientific interpretation of what these AI systems are actually doing.
Reasoning comes in many technically defined forms, but the basic procedure is easily recognizable: arriving at a sound conclusion by linking together intermediate steps that logically follow from each other. We do this with thoughts; LRMs use so-called chains of thought, a term of art for the streams of synthetic text that the models emit before arriving at an answer to a complex query. One minute, the idea that AI could reason via these chains was being prominently and credibly critiqued (by a team of researchers from Apple) as an “Illusion of Thinking” subject to “complete accuracy collapse” under surprisingly simple conditions. The next minute, LRMs were bagging gold medals at the International Mathematical Olympiad, a feat so challenging that “even very successful mathematicians and scientists may well highlight [it] on their CVs all their lives,” as the scientist and AI critic Gary Marcus and Ernest Davis wrote in 2025. If that’s not a sign of “real” reasoning, what is?
p]:my-6 [&>ul]:my-6 [&>ol]:my-6 [&>p]:text-3-5 [&>p]:leading-6.5 [&>li]:text-3-5 [&>li]:leading-6.5 [&_img.alignleft]:float-left [&_img.alignleft]:mr-5 [&_img.alignleft]:ml-0 [&_img.alignleft]:my-5 [&_img.alignright]:float-right [&_img.alignright]:ml-5 [&_img.alignright]:mr-0 [&_img.alignright]:my-5 [&_figure]:m-0 [&_figcaption]:relative [&_figcaption]:flex [&_figcaption]:flex-col [&_figcaption]:gap-2 [&_figcaption]:pt-2 [&_figcaption]:pb-4-5 [&_figcaption]:mt-0 [&_figcaption]:mb-6 &_figcaption]:font-pangram [&_figcaption]:after:content-[""] [&_figcaption]:after:absolute [&_figcaption]:after:bottom-0 [&_figcaption]:after:w-11 [&_figcaption]:after:h-0.5 [&_figcaption]:after:bg-gray-1a1 [&_.caption]:block [&_.caption]:font-pangram [&_.caption]:text-0xxs [&_.caption]:leading-4-5 [&_.caption]:m-0 [&_.attribution]:block [&_.attribution]:font-pangram [&_.attribution]:text-xs [&_.attribution]:leading-4-5 [&_.attribution]:m-0 [&_.attribution]:before:content-none show-dropcap" style="color: #000000;"><br>I n philosophy, “qualia” refers to the subjective qualities of our experience: what it’s like for Alice to see blue or for Bob to feel delighted. Qualia are “the ways things seem to us,” as the late philosopher Daniel Dennett put it. In these essays, our columnists follow their curiosity, and explore important but not necessarily answerable scientific questions.
But wait — soon after, more research, from the Santa Fe Institute, showed that LRMs can crush even carefully designed benchmarks for reasoning (like a collection of analogy-like visual puzzles) using mere “surface-level ‘shortcuts.’” What they were doing looked less like generalizable reasoning than just gaming the system. Then, as if on cue, another “hold my beer” moment: Google DeepMind and the mathematician Terence Tao (the GOAT!) used AI to rediscover or improve the solutions to 67 problems “spanning mathematical analysis, combinatorics, geometry, and number theory.” Deal with it, haters!
What about additional evidence that LRMs can’t reason reliably, even when they possess the necessary algorithm and computational budget to do so, and suffer from a list of scientifically documented failure states long enough to use as a Slip ’N Slide? Whatever — I guess that’s just “jagged intelligence” for you...