Anthropic says Claude formalized Fermat's Last Theorem in Lean in 11 days
- Anthropic says Claude produced the first complete computer-checked proof of Fermat's Last Theorem in Lean after working largely autonomously for 11 days.
- Claude wrote 13 million lines of Lean and proved 29,500 intermediate theorems while formalizing the result, according to Anthropic.
- The project verifies a known theorem rather than presenting new mathematics: Lean checks each logical step against its formal system, including steps that human-written proofs omit.
- Kevin Buzzard's Lean formalization project began as a community effort in 2024; Buzzard said Anthropic's repository proves the whole theorem, while his project continues work on Mathlib objects and human-readable exploration tools.
- Anthropic says easier formalization could reduce the time needed to check mathematical results, while the published repository makes the generated proof available for inspection.
Hacker News opinions
Buzzard's group got scooped, but he seems to be taking it well. His response calls it an extraordinary achievement and says the proof builds usable multi-layered formalization artifacts.
Buzzard says his EPSRC project still has work to do. Anthropic formalized the whole theorem, while his project is also adding modern number-theory objects to Mathlib and building a document for humans to explore the modern proof.
The 11-day figure is striking, though Buzzard's five-year project had £1M in funding and Anthropic may have spent far more.
Claude's result appears to tick the final item on Freek Wiedijk's list of 100 formally proved theorems.
I find Lean hard to read even after spending time on advanced mathematics and several Lean introductions. Coq and Isabelle feel closer to pen-and-paper proofs to me, and I wish proof languages were more digestible.
Lean is not necessarily the future of proofs for humans. There are many proof assistants, much as there are many programming languages, and a machine-checkable corpus should be translatable between them.
Thirteen million lines and 29,500 intermediate theorems make the result seem to support the idea that models can do work whose correctness can be mechanically shown.
The next step should be refactoring. This proof is far too large, and getting a concise version accepted into Mathlib would be a more meaningful finishing point.
I do not see why 13 million AI-generated Lean lines should automatically inspire confidence. Have we replaced checking a human proof with checking a vast generated codebase?
The repository labels the result "self-assessed," and Lean and Nanoda kernels have previously missed the Collatz hack. I would like an independent translation into HOL Light.
Moving from trust in a human proof to trust in Lean still raises confidence dramatically. A second prover might add confidence, but no kernel implementation is immune to bugs.
The relevant metric may be cost, not the claimed 11 days. We do not know how much compute a frontier lab used.