MathOverflow asks if AI compute swarms are dragging mathematics back into secrecy, as Terence Tao says finding a problem is now the scarce resource
- Terence Tao warned on Mastodon that the identification of a promising problem is now the scarce resource, since even a rumor of someone working on a problem triggers massive AI-powered effort to flatten it, and the incentives now point toward no longer sharing promising research directions with the community.
- Cost estimates for the AI effort behind the claimed Navier-Stokes counterexample range from about $6.5M on compute (Capital&Compute) to $10-15M (TensorFeed) and $10-40M in total costs (Business Insider), which the question argues only capital-rich companies can spend.
- The poster draws a parallel to medieval practice: Scipione del Ferro found a method for depressed cubics around 1510 but kept it hidden for 20 years and disclosed it only on his deathbed, and Newton's delayed publication of calculus fed the Royal Society dispute with Leibniz.
- LLMs now produce largely correct proofs for undergraduate, graduate, and previously unsolved problems, and can output a full Lean formalization, a situation one linked essay calls the fall of the proof economy.
- Commenters propose remedies such as university-hosted open-weight models, publication clauses restricting research use in AI training, and regulation of AI companies, while noting no individual can afford the infrastructure alone.
Hacker News opinions
Researchers need to stop sharing freely with AI companies and write explicit clauses into their publications about consent. There should be a clear legal distinction between using research data for model training and another researcher using it, with one scientific body enforcing it.
A clause in a publication won't stop it from being used as training data. Information wants to be free. The frontier vendors do sell enterprise licenses that contractually guarantee your prompts aren't used for training, so scholars who care about attribution will have to buy those or run their own open-weight instances.
Would that legal framework cut both ways? If an AI company uses AI to make a mathematical discovery and publishes it, could it then legally stop human mathematicians from using it?
Universities could just host open source models. They can afford it even when individual mathematicians can't, and if small math-optimized agents show up like the coding ones, those might run on a personal machine.
Simple answer: if you don't publish your work, we don't fund you. Why is this even a question?
That misses the point. This is about cooperation before you publish. They will keep everything medieval secret, otherwise some big company steals it and claims it as their own.
Stop treating AI companies and their software as unconstrained, above-the-law actors. Regulate them. At the same time mathematicians should be using specialized LLM tools in far more sophisticated ways than lay people; if you keep using your slide rule, you won't keep up.
Same logic as saying we can prevent the rest of society from devolving into secrecy: treat people with respect and dignity instead of as ore to be profitably extracted.
Any solution that isn't a socialist pipe dream? Telling people that is like telling depressed people to just stop being depressed.
We don't prevent it. Before LLMs it took some effort to snipe someone and your reputation was on the line. Now anyone with a few bucks can do it, same as YouTube AI slop: hum a few bars and 27 people have posted rip-offs within the hour.
Seeing the same thing in video games. Easy-to-make clones with generic names used to get copied once, now ten people clone them the same week.
Honestly I welcome the era of secrecy. After living in this open information age where everyone seems to know everything, secrets might make things interesting again. Maybe the wise ones go back to paper.
There's real money in selling access to the leading LLMs you need to find whatever secret result is out there. Same as the PlayStation hypervisor 0day: if an LLM can find someone's secret 0day, it can find a math proof someone says they have.
Two worries keep coming up: people will publish so much frontier math that humans can't understand it, and frontier math will all be kept secret. Those can't both happen at once.
The real worry is different. AI-generated frontier math makes it hard to identify who the actual frontier mathematicians are, and trained frontier mathematicians become scarce. The worse of both can be true at once.
Universities should host LLMs for their faculty and students, the way they run any other computer lab. Nobody should be using public LLMs for their work, it should be against the rules to submit someone else's work to a commercial LLM, and universities should offer models that aren't datamined.