The Crisis in Mathematics and the Prospect of AIcademia
MY SITE
Home
Academic Philosophy
Research
Teaching
Logic Textbooks
Research Groups
Private Tutoring
Popular Philosophy
Talking In Circles
Absolute Irony (blog)
To Big Finities and Beyond
Making Sense of It
Paintings
About Me
Absolute Irony
Philosophical musings on<br>whatever suits my fancy
A continuation of the old blog, found here
The Crisis in Mathematics and the Prospect of AIcademia
8/1/2026
3 Comments
Mathematics is in crisis. Or, at least, mathematicians are. Mathematics itself, on the other hand, might be entering a golden age. This strange tension is likely to confront a number of academic disciplines in the coming few years, and so, even academics who aren’t mathematicians themselves should probably be thinking about the situation currently facing mathematics, since something like it will may face them soon too. I am myself a philosopher rather than a mathematician. However, as a philosophical logician (among other things), I am more mathematician-adjacent than most of my colleagues whose work fits more squarely in the humanities. So I have been impacted by the crisis in mathematics more than most in my field, and it has led me to think about the future of academic research, especially in the sciences. The conclusion I’ve come to is quite a humbling one, at least for us humans.
The Crisis in Mathematics
When LLMs like ChatGPT first burst onto the scene just a few years ago, many were impressed with their wide range of linguistic capabilities, for instance, their poetry writing abilities. This was, in some sense, unsurprising; they were, after all, language models. While the linguistic abilities of these systems were impressive, they were widely mocked for their utter mathematical incompetence. Mathematicians, it seemed, were in the clear. However, since the release of “reasoning models,” first with OpenAI’s “o1” in September 2024, then with “o3” in April 2025, the writing has been on the wall that these systems were coming for mathematics.
These new “reasoning models” are trained through large-scale reinforcement learning to engage in an extended internal “chain of thought” before producing a final answer. That is, they are trained to break problems into steps, recognize and correct mistakes, and abandon unsuccessful approaches for new ones, doing all of this “internally” before they submit a final answer to the user. Their performance can then be improved by scaling both the training compute used to reinforce successful reasoning behavior of this sort and, crucially, the amount of inference-time compute they are permitted to spend working through a problem. These models quickly became very very good at tasks in which success could be verified, most notably, coding and math.
Just over a year ago, both OpenAI and Google announced their models achieving gold medal performance in the International Mathematical Olympiad, a set of competition problems designed to challenge the most mathematically-gifted high school students. Since then, model capabilities have progressed beyond self-contained problems to genuine research mathematics. Over the last few months, a number of notable conjectures—most notably, the Unit Distance Conjecture and the Jacobian Conjecture (for dimensions greater than 2)—which had stumped mathematicians for decades have been solved by large language models, the former by an internal model of ChatGPT and the latter by Claude Fable 5. I will not go into the details of what these mathematical conjectures say, but I will note that they were very significant open problems in their respective fields.
In both of the two cases just mentioned, the LLM did not prove the conjecture. Rather, it produced a counterexample, disproving the conjecture. This has been the general pattern of the most prominent results in the last few weeks since the newest class of models have been made public. Each day now, it seems, more and more conjectures are falling at the hands of LLMs. The prompts for some of these results are quite comical, and, to mathematicians, I’m sure depressing.
Last week, Dmitry Rybin posted a ChatGPT-generated counterexample the Dinitz-Garg-Goemans conjecture along with the prompts he used to get ChatGPT to generate it, the first of which included the simple command “You should do a breakthrough.” When it came back from its attempts an hour later with no such breakthrough, Rybin simply urged it “Continue the search. Have a clear strategy obtained from deeper understanding of the problem structure.” Ninety minutes later, still no conclusive counterexample, only a partial result that did not suffice to refute the conjecture. Rybin urged it again: “enough of partial results. Let’s finish with a complete unconditional counterexample.” Ninety minutes later, it came back...