Assessment of open AI math results

paulpauper4 pts0 comments

Igor Kotenkov on X: "It's hard for an ordinary person to understand the complexity of these tasks. I'm no mathematician, and I don't see a difference between e.g., results 3 and 10.

So I had GPT-5.6 Sol Pro and Fable 5 Max classify these using @EpochAIResearch OpenMath's rubric:

— "Solid Result": A strong researcher in the area would be happy if their median output addressed problems of this caliber. Still, the problem would probably not get much engagement outside of its subfield.

— "Major Advance": The median person working in a broad area of mathematics (on the scale of number theory or graph theory) would take note, and would likely make the time to understand at least the outline of the solution.

— "Breakthrough": The median mathematician would want to know about this result, even if it was outside their area. It would be a candidate for one of the best results of the year in all of mathematics

Both Fable and Sol agree #3 is a Breakthrough (which explains why @SebastienBubeck opens his tweet with it). They also agree that at least 7 are Major Advancements. Fable thinks #7 is just a Solid Result, while Sol assigns the "Major Advance" label.

What's also interesting is that the official https://t.co/XdrU3D1VVf rubrics say this:<br>> When multiple tiers seemed plausible for a problem, we erred in the conservative direction. It would be disappointing to downgrade a problem’s notability after it was solved, whereas we can always highlight any unexpectedly interesting elements of a solution.

And Fable 5 thinks that at least 3 of the results are "Borderline Breakthrough" (#1, #4, and #9). @AcerFur any thoughts on this" / X<br>Post

Log inSign up

Post

Igor Kotenkov

@stalkermustang

It's hard for an ordinary person to understand the complexity of these tasks. I'm no mathematician, and I don't see a difference between e.g., results 3 and 10.

So I had GPT-5.6 Sol Pro and Fable 5 Max classify these using @EpochAIResearch OpenMath's rubric:

— "Solid Result": A strong researcher in the area would be happy if their median output addressed problems of this caliber. Still, the problem would probably not get much engagement outside of its subfield.

— "Major Advance": The median person working in a broad area of mathematics (on the scale of number theory or graph theory) would take note, and would likely make the time to understand at least the outline of the solution.

— "Breakthrough": The median mathematician would want to know about this result, even if it was outside their area. It would be a candidate for one of the best results of the year in all of mathematics

Both Fable and Sol agree #3 is a Breakthrough (which explains why @SebastienBubeck opens his tweet with it). They also agree that at least 7 are Major Advancements. Fable thinks #7 is just a Solid Result, while Sol assigns the "Major Advance" label.

What's also interesting is that the official Epoch.AI rubrics say this:<br>> When multiple tiers seemed plausible for a problem, we erred in the conservative direction. It would be disappointing to downgrade a problem’s notability after it was solved, whereas we can always highlight any unexpectedly interesting elements of a solution.

And Fable 5 thinks that at least 3 of the results are "Borderline Breakthrough" (#1, #4, and #9). @AcerFur any thoughts on this

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Noam Brown

@polynoamial

9h

An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.

We believe it will be a major step for scientific reasoning. openai.com/index/ten-adva…

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0">9:30 AM · Aug 1, 202614.5KViews

13<br>119<br>41

span:not(:empty)~span:not(:empty)]:before:content-['·'] [&>span:not(:empty)~span:not(:empty)]:before:px-1 [&>span:not(:empty)~span:not(:empty)]:before:shrink-0 min-w-0 overflow-hidden">Sean X. Luo MD PhD

@seanluomdphd

5h

I'm asking AI to explain to me about #3, and it does seem very important and potentially would go into an undergrad 1st year algebra textbook.

Intro algebra seems to avoid the most beautiful recent theorems on countable/finite groups as they involve confusing machinery.

svg]:size-5 text-body hover:bg-mix-current hover:bg-mix-amount-10 active:bg-mix-current active:bg-mix-amount-15 focus-visible:bg-mix-current focus-visible:bg-mix-amount-10 outline-current -m-2 shrink-0 cursor-pointer border-transparent p-0 text-body size-9 [&>svg]:size-[1.25em] [&>[data-engagement-icon]]:size-[1.25em] group-hover:bg-mix-current group-hover:bg-mix-amount-10" aria-label="View count" type="button" data-state="closed" href="/seanluomdphd/status/2083517909620367620/quotes">799

Log in or...

span empty major before fable results

Related Articles