🧮 On September 11, 2026, twenty-five Fields Medalists signed a joint declaration about AI and mathematics. Headlines around the world had them angry at AI for solving their problems. Read the actual text and something else is going on. The signatories open by conceding that AI has genuinely become good at mathematics. What they attack is not the technology. It is the yardstick.

What happened on September 11

The declaration went up on September 11 on the blog of mathematician Terence Tao, with the full text hosted at mathandai.org. Its title is "A Severe Misalignment of AI in Mathematics." Nikkei, Yomiuri and Asahi all reported the declaration without printing that title.

Twenty-five Fields Medalists signed it. The Fields Medal goes to mathematicians under forty, once every four years, and is often described as mathematics' equivalent of a Nobel. The signatories span forty-eight years of the prize, from Pierre Deligne, who won in 1978, to Yu Deng, who won in July 2026. Japan's Shigefumi Mori, the 1990 laureate, is on the list.

Tao explained that the text came out of roughly a week of discussion among mathematicians. They skipped the long consultative drafting behind comparable documents like the Leiden declaration, because they judged the situation urgent. No single author is named. It goes out under all twenty-five names, and as of September 16 it is still open: anyone can add an endorsement through ORCID or an academic email address.

The trigger came three days earlier. On September 8, OpenAI announced a result on the Navier-Stokes equations, which sit among the Millennium Prize Problems.

The target is not "solving"

The declaration begins with a concession that sits oddly with the coverage it received. Large language models, it says, have improved so sharply over the past few months that they now settle long-standing open problems across a range of mathematical fields. The signatories are not disputing the capability. That is where the argument starts.

What they call harmful is one specific practice: AI companies setting their models loose on mathematical problems in order to benchmark themselves against each other. A benchmark, in this context, is a shared test used to compare one model with another. That practice, the declaration says, is "detrimental to the science of mathematics, and to the mathematical community."

The subject of that sentence is AI companies, not AI. The object is using problems as a scoreboard, not solving them. The target is a metric, not a technology.

It would be strange otherwise. Tao is a well-known advocate of using AI in mathematics. By his own published record he has spent years promoting formalization in the Lean proof assistant, and he has posted sessions in which he used LLMs to check algebra and surface literature for his own research. He is not asking AI to leave the building.

Fast correct answers are not the same as understanding

The misalignment in the title is a mismatch of objective functions. What AI companies optimize for is producing true statements faster and in greater volume. What mathematics is for, the declaration argues, is conceptual understanding and insight, and passing both to the next generation. Solving a problem is a tool for getting there. It is not the destination.

The concrete grievances are about process rather than content. Results get pushed out too fast for anyone to write them up properly, to pull out whatever is new in the method so others can reuse it, or to credit the earlier work the result stands on. From there the declaration raises what it calls severe attribution and plagiarism questions. If nobody records whose idea did the work, the community cannot graft the result onto the existing body of mathematics.

Statements that can be marked true or false, mass-produced at an accelerating pace, may destroy the soil in which new ideas grow rather than fertilize it.

The proof that triggered it had exactly this problem

What OpenAI announced on September 8 was narrower than the coverage suggested. The company said it had established statements (C) and (D) as the problem is officially formulated: that under external forcing, an initially smooth three-dimensional flow can develop a singularity in finite time. The version most readers have in mind under that name is the unforced one, statements (A) and (B), and it is a different problem. OpenAI described running roughly ten thousand agents concurrently and reaching the result in about eighty-eight hours. Formalization in Lean took seventeen hours. The paper runs to 166 pages.

OpenAI also ruled out the money: "We do not intend to claim the Millennium Prize for this result." The Millennium Prize Problems were set by the Clay Mathematics Institute in Paris on May 24, 2000. There are seven of them, one million dollars each. According to reporting, Clay's president Martin Bridson said the institute's assessment would be deliberately unhurried and rigorous. Under Clay's rules a proof only becomes eligible once it has been published in a suitable venue, two years have passed, and the mathematical community has broadly accepted it. As of September 16 the institute's own page still lists Navier-Stokes as "Active."

Tristan Buckmaster of NYU and Levent Alpöge had posted their own result on the forced Euler equations roughly twelve hours before OpenAI went public. The two argued their unpublished work may have reached OpenAI; OpenAI denied any prior access, while conceding it could not exclude the possibility that anonymised usage data from its products had fed into its models. According to reporting, the initial version of OpenAI's paper cited neither Diego Córdoba and Luis Martínez-Zoroa, whose earlier techniques underpinned the whole approach, nor Buckmaster and Alpöge. The Córdoba citations were added in a revision the same day.

The declaration's abstractions (rushed announcements, missing citations, unclear attribution) were not hypothetical. They had just happened, three days before it was published.

And then the headlines dropped "AI companies"

In Japan, Nikkei ran "Math problems solved by AI: 25 Fields Medalists issue a condemnation, 'harmful to science.'" Yomiuri ran "Using AI to solve hard math problems is 'harmful to science and to mathematics': 25 mathematicians worldwide issue an urgent condemnation." In both, the words "AI companies" vanish, and what is being condemned becomes the act of solving. Asahi kept the subject, running "AI companies proving hard math problems is harmful," but landed on a headline that still reads as though proofs themselves are the problem.

The bodies told a different story. Asahi described the target accurately as companies having AI solve mathematical problems as an indicator of AI performance, and printed the mathandai.org URL. What broke was the headline layer.

Nor did every Japanese outlet get it wrong. GIGAZINE's headline named the actual dispute: mathematicians are pushing back on the use of mathematics as an AI benchmark. THE BRIDGE linked both to Tao's blog and to mathandai.org and noted that endorsements are still being collected.

The same compression happened in English. "Rapid AI Proofs Are Harming Math" and "AI Problem-Solving Competitions Are Eroding the Foundations of Mathematics" sat alongside more careful framings such as "25 Fields Medalists Warn AI Math Benchmarks Erode Attribution and Auditability." This is not a Japanese failure. It is a headline failure, and it happened in at least two languages at once.

One of the Japanese mathematicians who signed the declaration wrote on his own blog that he was astonished to see the Nikkei and Yomiuri headlines drop the phrase "AI companies."

Why is this statement so hard to headline? Because the thing being criticized is neither a technology nor a product but an evaluation metric, which is an abstraction. Headlines want a subject and a verb. "Mathematicians versus AI" lands instantly. "Mathematicians versus how progress is measured" does not. When you compress, that distinction is the first thing to go.

This is not only about mathematics

The declaration widens its own scope in its closing lines. The problems mathematics faces now, it says, resemble those facing other scientific and creative professions, and point to problems all of humanity may face. It ends by asking that all of this be taken up urgently: by mathematicians themselves, by the firms building the technology, and by society at large.

Benchmarks are everywhere outside mathematics. Pass rates on coding tests. Automated translation scores. Performance on medical licensing exams. Each is a convenient proxy for progress, and the more convenient it gets, the more it crowds out whatever the field was actually trying to do. Mathematics may simply have gone first, because its answers can be checked by machine.

Endorsements were still being collected as of September 16, and the argument is not over. So: what is your work measured by right now, and does that yardstick point in the same direction as the thing you are actually trying to do? How did this story get reported where you live?

References