OpenAI releases frontier model math results on GitHub with Lean proofs and reasoning traces
- OpenAI published a batch of new mathematical results produced by an internal frontier model in a public GitHub repository (github.com/openai/math), with protocols for paper revisions and citations, after consulting the Advisory Group on Mathematics and AI at the Institute for Advanced Study.
- The release ships Lean formalizations of many proofs, 10 summaries of the model's reasoning, compute estimates expressed as ChatGPT Pro usage, and statistics on attempted problems; OpenAI says the average result took roughly three hours of ChatGPT Pro thinking.
- Hacker News readers single out the Unique Games Conjecture work as the biggest item, calling it more consequential than a Millennium Prize problem, alongside claimed progress on Riemann and Hodge related questions.
- OpenAI says it will fund workshops, conferences, and special programs on understanding major AI-produced results, and is working to responsibly release the model that produced them.
- OpenAI commits to better citations, mathematical exposition, and presentation in future releases, and says it is still exploring community-hosted alternatives that meet the advisory group's guidelines.
Hacker News opinions
Good to see OpenAI actually engaging with the math community, even if they had to be publicly shamed into doing it.
Honestly this is the opposite of what the math community was asking for, and I say that as someone who strongly approves of the approach.
It's wild that gatekeeping got to the point where they felt they needed permission to share math. They aren't asking for permission, they're framing it that way because of the bad press.
I think it's fine. OpenAI noticing when a project steps on a human community and then respecting their norms is a good thing, especially when the results have no immediate application and build on thousands of years of that community's work.
After how badly OpenAI and Anthropic handled the earlier cases, this more cautious rollout is an improvement. If they want to publish in rigorous fields they have to match the rigor, not drag it down to the level of ML conference papers.
So happy this is on GitHub instead of some paywalled journal. Truly a new age for science.
Most math results go to arXiv, and journals add peer review on top. GitHub is a much worse place to store important results, and slop shouldn't end up on arXiv either.
I don't think this hits OpenAI where it needed to hit if this is their primary response.
You want them to stop doing math research?
Can someone who actually knows the field outline which of these results matter most? I can't tell what's load bearing from the announcement.
The reasoning traces are the fun part. One of them has a line about a cheater choosing arbitrary g_{v_i} and passing on duplicated input, and the model is clearly delighted with itself. It's also only an excerpt, so who knows what got left out.
The actual results are in overview.pdf in the repo, and CONTENTS.md is the HTML version if you don't want to open a PDF.
This is significant progress and it landed without all the drama. Riemann, Hodge, Unique Games, all in one drop. Point your agent at the repo and ask it what the significance is. Feels like 50 to 100 years of human math progress.
Not to pick on you, but the number of supposed math enthusiasts who can't spell Riemann is genuinely shocking. Also there was A LOT of drama around this release.
Unique Games is a foundational conjecture in complexity theory and an assumption behind a huge number of inapproximability results, so a valid proof is a big deal. To me that's bigger than a Millennium Prize problem.
I'm glad they're formalizing in Lean, because I'm not convinced these models can write down their thoughts in English. Section 1.1 of the Unique Games paper spends two sentences on elementary graph notation and it reads badly.
Hilarious to read this right next to AGMAI's requests. Basically: here you go, have fun, ignore all your demands, and by the way we're releasing the model, stay tuned.
Hypothetically I'm a PhD student halfway through and one of these preprints overlaps my thesis. Do I pivot the whole thing? Don't paste your research into these models, they'll train on it and scoop you. Treat what's published as the new base and go forward.
People get scooped by other humans constantly, that's just research. It stings, but it's a signal you were working on something other people care about.
My entire PhD got made obsolete by CRISPR and it was wonderful. Look forward to being obsolete.