OpenAI withdraws three math manuscripts from its GitHub repo as Lean checks continue
- OpenAI's update to its GitHub math repo added 6 new Lean formalizations, made 19 modifications, and withdrew 3 results, leaving about 42% of top-line results formally verified.
- The withdrawn manuscripts are Algebraicity of Weil classes on split abelian eightfolds, Algebraicity of Kuga-Satake Correspondences for K3 Surfaces, and The rational Hodge conjecture for products of K3 surfaces.
- The vast majority of results came from an unreleased internal OpenAI model, with about 3 hours of ChatGPT Pro thinking compute per result across roughly 4,000 posed problems.
- OpenAI's original post said Lean formalizations would be added later, so some results were published before every proof was checked.
Hacker News opinions
Early write-ups read like slop to me. My guess is nobody at OpenAI can validate the output well enough, so whatever reads well gets published.
Lean checks the proof, not whether the statement is what you meant. A formalization can be semantically off for a slightly different problem and still compile, and that's the real gap.
Their original blog said they'd add Lean formalizations later, so they published first and checked later. The 42% number is them catching up.
Lean definitions aren't a black art. Do the tutorial and you can read the statement of most of these results. Following the proofs is the hard part.
Fermat's Last Theorem had a massive flaw that took two years and outside help to fix, and the abc conjecture proof is widely believed false. Human proofs aren't clean either.
The matrix multiplication bound is theoretically faster but only slightly better than the previous 2.371177 at real sizes. It's a galactic algorithm.
The repo says the internal model was posed about 4,000 problems and only around 700 got posted. Seeing the failures would tell us a lot about what's hard for LLMs.
An IPO next year means mathematicians have a few months to find errors before then. If there are real errors and nobody catches them, I'd have questions about the math community.