Is AI Reasoning Right for the Wrong Reasons?

retupmoc011 pts0 comments

Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazine

About Quanta

Search

Search for:

Search<br>Search

Newsletter

Get the latest news delivered to your inbox.

Email

Subscribe

Recent newsletters

Follow Quanta

Facebook

Youtube

Instagram

RSS

An editorially independent publication supported by the Simons Foundation.

Type search term(s) and press enter

What are you looking for?

Search

Home

Is AI Reasoning Right for the Wrong Reasons?

Comment

Save Article

Read Later

Share

Facebook

Copied!

Copy link

Email

Pocket

Reddit

Ycombinator

Comment

Comments

Save Article<br>Read Later

Read Later

Qualia

Is AI Reasoning Right for the Wrong Reasons?

By

John Pavlus

July 31, 2026

The idea that artificial intelligence can “reason” is more intuitive than ever. But intuitions can be wrong, and the science is far from settled.

Comment

Save Article

Read Later

Celsius Pictor for Quanta Magazine

Introduction

By John Pavlus

Contributing Writer

July 31, 2026

View PDF/Print Mode

artificial intelligence

computer science

features

large language models

machine learning

Qualia

All topics

I’ll just say it: What the hell is going on with AI “reasoning”?

Sorry for the air quotes. That punctuational side-eye was more common in 2024, when the specially trained cousins of LLMs now known as “large reasoning models,” or LRMs, were still new. Nowadays it may seem downright churlish, though, given that a “general-purpose reasoning model” from OpenAI solved a famous open mathematical research problem in one shot in May 2026. Still, I’m not sure how else to acknowledge my intellectual whiplash over the scientific interpretation of what these AI systems are actually doing.

Reasoning comes in many technically defined forms, but the basic procedure is easily recognizable: arriving at a sound conclusion by linking together intermediate steps that logically follow from each other. We do this with thoughts; LRMs use so-called chains of thought, a term of art for the streams of synthetic text that the models emit before arriving at an answer to a complex query. One minute, the idea that AI could reason via these chains was being prominently and credibly critiqued (by a team of researchers from Apple) as an “Illusion of Thinking” subject to “complete accuracy collapse” under surprisingly simple conditions. The next minute, LRMs were bagging gold medals at the International Mathematical Olympiad, a feat so challenging that “even very successful mathematicians and scientists may well highlight [it] on their CVs all their lives,” as the scientist and AI critic Gary Marcus and Ernest Davis wrote in 2025. If that’s not a sign of “real” reasoning, what is?

p]:my-6 [&>ul]:my-6 [&>ol]:my-6 [&>p]:text-3-5 [&>p]:leading-6.5 [&>li]:text-3-5 [&>li]:leading-6.5 [&_img.alignleft]:float-left [&_img.alignleft]:mr-5 [&_img.alignleft]:ml-0 [&_img.alignleft]:my-5 [&_img.alignright]:float-right [&_img.alignright]:ml-5 [&_img.alignright]:mr-0 [&_img.alignright]:my-5 [&_figure]:m-0 [&_figcaption]:relative [&_figcaption]:flex [&_figcaption]:flex-col [&_figcaption]:gap-2 [&_figcaption]:pt-2 [&_figcaption]:pb-4-5 [&_figcaption]:mt-0 [&_figcaption]:mb-6 &_figcaption]:font-pangram [&_figcaption]:after:content-[""] [&_figcaption]:after:absolute [&_figcaption]:after:bottom-0 [&_figcaption]:after:w-11 [&_figcaption]:after:h-0.5 [&_figcaption]:after:bg-gray-1a1 [&_.caption]:block [&_.caption]:font-pangram [&_.caption]:text-0xxs [&_.caption]:leading-4-5 [&_.caption]:m-0 [&_.attribution]:block [&_.attribution]:font-pangram [&_.attribution]:text-xs [&_.attribution]:leading-4-5 [&_.attribution]:m-0 [&_.attribution]:before:content-none show-dropcap" style="color: #000000;"><br>I n philosophy, “qualia” refers to the subjective qualities of our experience: what it’s like for Alice to see blue or for Bob to feel delighted. Qualia are “the ways things seem to us,” as the late philosopher Daniel Dennett put it. In these essays, our columnists follow their curiosity, and explore important but not necessarily answerable scientific questions.

But wait — soon after, more research, from the Santa Fe Institute, showed that LRMs can crush even carefully designed benchmarks for reasoning (like a collection of analogy-like visual puzzles) using mere “surface-level ‘shortcuts.’” What they were doing looked less like generalizable reasoning than just gaming the system. Then, as if on cue, another “hold my beer” moment: Google DeepMind and the mathematician Terence Tao (the GOAT!) used AI to rediscover or improve the solutions to 67 problems “spanning mathematical analysis, combinatorics, geometry, and number theory.” Deal with it, haters!

What about additional evidence that LRMs can’t reason reliably, even when they possess the necessary algorithm and computational budget to do so, and suffer from a list of scientifically documented failure states long enough to use as a Slip ’N Slide? Whatever — I guess that’s just “jagged intelligence” for you...

_figcaption reasoning _img after search from

Related Articles