mrkeyoor.com_
Sat 12 Sept 03:57 UTC
AI7 min read

25 Fields Medalists Say AI Proofs Are Creating Review Debt

A new declaration says AI labs can generate mathematical claims faster than researchers can absorb them. Its practical target is the publication and review pipeline.

By 19:30 UTC on September 11, a Hacker News submission about a protest from leading mathematicians had reached 624 points and 675 comments, less than two hours after it was posted. The count measures attention alone; it cannot verify the declaration's claims. What it does reveal is how readily developers recognized the failure mode behind the complaint: machine output is getting cheaper to produce while expert review remains slow and scarce, with people still accountable for it.

The declaration carries the names of 25 Fields Medal recipients, including Terence Tao, Peter Scholze, Maryna Viazovska and James Maynard. They argue that AI companies are treating famous unsolved problems as model benchmarks, then leaving mathematicians to check the work and reconstruct how its ideas connect to earlier research. In software terms, the model can open pull requests at machine speed while the maintainers still own the merge.

What the signers object to

Famous problems have traditionally done more than furnish a correct-or-incorrect score. The signers describe them as reference points around which researchers develop methods, teach students and turn a difficult proof into an explanation other people can use. A solution matters partly because of the discussion and rewriting that follow it. A bare answer captures only the endpoint of that process.

That distinction changes the economics of automated mathematics. A lab with enough compute can search many possible arguments at once and announce a candidate result as soon as one survives its internal checks. The declaration says the resulting flood of true-or-false statements could consume the very problems that help young researchers acquire judgment. Each candidate may also require specialists to identify prior art and decide whether the argument contains a reusable idea. The output arrives quickly; the field inherits the queue.

The signers focus on rushed publication as much as raw capability. They say some AI-produced solutions have been announced before a proper write-up, an account of the new methods or adequate citations to earlier work. That creates concrete questions about attribution and plagiarism. It also asks mathematicians to perform the unpaid work of translating a machine result into the shared body of knowledge, according to the declaration's account of the research process.

This is not a call to ban mathematical AI. The same declaration says the technology could accelerate genuine study and that the profession will have to adapt. Its dividing line is whether automated problem solving feeds human understanding or substitutes a stream of claims for it. The authors place responsibility for that choice on the people and organizations controlling the systems.

A one-week response to a longer dispute

Tao says the statement grew out of discussions among the 25 initial signatories over a single week. In his announcement, he also acknowledges that the group did not have time for a wider consultation. They chose a fast release because they considered the situation urgent. That compressed drafting process makes the document read like an alarm rather than a policy manual.

The timing followed a public fight over authorship, private model use and the announcement of an AI-generated proof. TechCrunch reported that New York University mathematician Tristan Buckmaster accused OpenAI of pressuring him to omit credit for a collaborator employed by Anthropic. The outlet also reported that OpenAI withdrew sponsorship from a Caltech mathematics event after criticism from researchers there. Those are attributed allegations and actions surrounding the release, rather than proof of the declaration's broader case.

The declaration itself never names OpenAI, Buckmaster or a particular theorem. Its argument applies to any lab that turns an open problem into a product demonstration before the mathematical community has had time to inspect and absorb the result. That scope keeps the issue alive even if the immediate authorship dispute is resolved. Publication incentives and review capacity remain.

Review is the scarce resource

Mathematicians had already written a more detailed rulebook. The Leiden Declaration on Artificial Intelligence and Mathematics, dated June 2 and endorsed by the International Mathematical Union, was developed over eight months after a 2025 conference. It asks researchers to disclose automated tools and computational resources, provide precise references, and keep human authors responsible for correctness. It also asks professional organizations to prepare for major results produced by unconventional means.

Leiden's proposed review safeguards are specific. Depending on the work, an organization could require a human account of the central argument, formal verification, cross-checks between theoretical and computational results, or external review before submission. It also says press releases and blog posts can support publication but cannot replace peer review and community scrutiny. These rules attach a review budget to the claim instead of assuming the field will supply one afterward.

A formal proof assistant helps with part of that job. The Lean language reference says its editor's check marks mean the kernel accepted a proof of the stated theorem from the definitions, axioms and imports in the project. The same documentation separates two questions: whether a theorem has a valid formal proof, and what its statement means. Trust still depends on the formal statement matching the intended informal claim and on the imported material avoiding unsound assumptions.

That boundary is familiar to anyone who has reviewed generated code. A passing test suite can establish defined behaviors while missing a wrong requirement, an unsafe dependency or an architectural mistake. Formal checking gives mathematics a stronger instrument than ordinary software tests, but the translation into definitions and the relevance of the result still demand human judgment. Both the Lean documentation and Leiden declaration make that division explicit.

When a benchmark changes the field

Mathematics is attractive to AI labs because many answers can be checked. The Leiden document says formalized proofs can provide a large supply of automatically verified feedback for training, while developers hope theorem-proving ability will transfer to broader reasoning. Famous open problems add a public scoreboard. A model that appears to solve one supplies a cleaner marketing story than a gradual improvement in proof search.

The new declaration argues that this scoreboard can alter the activity it measures. Researchers often share partial ideas in talks and private discussions before a paper is ready. Once a hint can trigger a well-funded automated search, early disclosure carries a new risk: another organization may reach and announce the endpoint first. TechCrunch describes mathematicians worrying that this pressure will encourage secrecy, a concern that follows from the current dispute rather than an established change in behavior.

Less sharing would also reduce the human material from which future systems and researchers learn. The 25 signers describe students and ideas as the profession's most precious resources, developed through discussion and careful writing over time. Their objection is therefore partly about a feedback loop: using the mathematical commons to accelerate results may weaken the practices that replenish that commons. Whether that happens will depend on how labs handle provenance, credit and release timing.

The practical publication contract

The new statement ends with an urgent request for action but offers no required release format. Leiden supplies the operational pieces. Researchers disclose the systems and computational resources they used. Human authors accept responsibility for the theorem and its citations. Publishers can request an accessible account of the argument, formal artifacts where appropriate and outside review before a public claim receives institutional weight. The full recommendations also ask research organizations to give academics legal support when they collaborate with industry.

For developers building mathematical agents, those proposals amount to a release contract. Publish the exact theorem statement and its dependencies. Put the machine-generated artifact beside a human explanation and a provenance record. State who checked it, what those checks establish and which parts remain unreviewed. A green proof-checker result can then be read at its proper scope instead of being stretched into a verdict on meaning or novelty, much less credit, as the Lean validation guide cautions.

The Hacker News total shows that this problem has escaped a specialist seminar, but the 624 points do not settle any mathematical or ethical question. The stronger evidence is the agreement between a fast protest signed by 25 Fields Medal recipients and the earlier Leiden process: both place responsibility on the people who publish automated results. The fast response adds urgency, while the earlier document gives publishers and research groups rules they can inspect.

What to watch in the next release

Concrete follow-through will be visible in publication practice. Watch whether journals adopt tool-and-compute disclosures, whether AI labs fund independent review, and whether the Math and AI declaration attracts endorsements beyond its original 25 signers. Leiden already records the International Mathematical Union's endorsement and calls for professional bodies to prepare for unconventional claims.

The next major AI-generated proof will provide the cleanest test. Its release should arrive with an exact statement, a traceable account of prior work, human-readable reasoning, formal artifacts where they help and an honest review status. If those pieces appear weeks later, or never, the speed celebrated in the announcement will still be transferring the hard part to outside mathematicians, the cost identified by both the new statement and the Leiden recommendations.

We reviewed this

  1. learn — our honest review
  2. requests — our honest review
  3. paper — our honest review

Sources

  1. Hacker News: A misalignment of AI in mathematics
  2. A Severe Misalignment of AI in Mathematics
  3. Terence Tao: A Severe Misalignment of AI in Mathematics
  4. OpenAI's feud with mathematicians is only escalating
  5. Leiden Declaration on Artificial Intelligence and Mathematics
  6. Lean Language Reference: Validating a Lean Proof