A release that left the checking to mathematicians
The AHM statement, published as a guest post on Terence Tao’s blog, says OpenAI disregarded a recommendation from the advisory group AGMAI: difficult mathematical problems should not be tested on internal models. The association says mathematicians did not ask for this work and calls the simultaneous release of more than 700 files a display of power, not scientific research.
The statement also points to the many copyright lawsuits OpenAI faces worldwide. It stops short of claiming the proofs came from training on mathematicians’ work, but raises the possibility that they did.
That concern sits alongside a more immediate dispute. OpenAI previously said an internal model had solved more than 100 open mathematical problems in a month, including the Navier–Stokes problem, one of the Millennium Prize Problems. The claim came shortly before two mathematicians planned to present their own AI-assisted solution to Navier–Stokes using ChatGPT. OpenAI denied suspicions that it had trained on those interactions and gained an advantage.
The release model matters, too. Scott Aaronson contrasts OpenAI’s batch of raw proof drafts with what he calls Anthropic’s approach. Anthropic worked with two algorithm researchers; its model suggested a key idea that helped disprove two conjectures that had stood for about twenty years. The researchers were paid and prepared an account of the proof that people could follow.
Neither approach resolves the underlying problem:
AGMAI, formed at the Institute for Advanced Study after criticism of OpenAI’s earlier claims, has already said it cannot influence the pace of the company’s internal research. That limit makes the advisory process look less like oversight than commentary after the fact.
The cost is not just verification
Tao has argued that mathematics is about more than solving problems. A statement he signed with 24 other Fields Medalists, including Peter Scholze, Marina Vyazovskaya and Martin Hairer, warned of a “serious divergence” between the aims of the AI industry and those of mathematics. Problem-solving, they say, is a means to conceptual understanding. Producing answers at scale and ever greater speed could “destroy fertile ground instead of bringing new ideas to life.”
The same statement criticizes rushed AI-result announcements that lack full explanations and references to important earlier work, raising questions about attribution and plagiarism.
Tao’s account of “Mathematics 2.0” points to a social cost as well. In the traditional pattern, solving a longstanding problem led to talks, seminars, collaborations and eventually textbooks. Young researchers helped interpret the result and place it in the discipline. Now, he says, AI may solve problems while the people who run the models are not interested in the mathematics or equipped to explain it. There are fewer seminars and collaborations, and researchers may hide promising directions for fear that competitors will reach an answer first.
For Tao, a result can also close off inquiry. Once a problem is declared solved, it cannot simply become open again; even knowing that a solution exists can “contaminate” the search for other approaches. He has proposed measuring progress not only by solved problems, but also by the quality of explanations, the health of the mathematical community and the creation of new research directions.
That proposal extends his 2024 idea of “industrial mathematics,” when people still set the pace and AI was meant to help as engines help chess players: search broadly but relatively shallowly, complementing experts’ deeper work. The current pattern has reversed. I think that reversal is the central tension here: the field is being asked to judge results produced faster than its institutions can interpret them.
Three hours against a career
Aaronson makes the imbalance personal through the experience of his wife, Dana Moshkovitz, who has spent her career on the Unique Games conjecture, an open problem in theoretical computer science. A proof of the conjecture appeared in the released files.
On the night of publication, Moshkovitz messaged Aaronson that the text resembled the work of someone on psychedelics. Much of it was incomprehensible. It cited many papers without explaining why they applied, even though earlier results should have ruled out the approach. She said the paper was so poorly written that it could not be read without AI help, and described its central construction as “some kind of alien nonsense.”
Aaronson says the collection also includes results in complexity theory, number theory and algorithm design, including partial progress on several remaining Clay Institute Millennium Problems. In his view, each could have been among the year’s major achievements on its own.
The set does not include cryptography. Citing unnamed sources, Aaronson says AI companies are now secretly searching for weaknesses in cryptographic protocols and basic mechanisms. About 8,000 problems were tested; roughly five percent produced successful results. Each took an average of three hours of GPT-Pro-level computation. He says the work probably also used OpenAI’s current internal model, which may reach paying ChatGPT users in the coming months.
I think the contrast between three hours of compute and a researcher’s career is the detail that makes this more than a dispute over bad writing. The release compresses the search for results, but it does not compress the work of deciding what those results mean.
A boycott cannot settle the question
Reactions to Tao’s post show a divided mathematical community. Some support a boycott and urge researchers to stop using OpenAI products. Others want better protection for preprint servers such as arXiv against large-scale data collection for training. Critics of the AHM position argue that OpenAI will not stop doing mathematics and that public results are preferable to secret ones. One commenter asks whether a problem can count as solved if neither its authors nor anyone else fully understands the proof. For Navier–Stokes, it remains unclear whether the result deepened anyone’s understanding of fluid dynamics.
AGMAI’s position is more cautious than AHM’s: publication is a first step, it says, and the work of understanding the results is only beginning. But it also warns that mathematical research cannot become a process of analyzing results produced by AI labs. Mathematicians need room to choose questions, develop methods and pursue directions that were not selected as demonstrations of a system’s capabilities. That freedom, the group says, requires equal access to powerful research tools and sufficient computing resources.
What I’d want to know is whether those conditions can be met while private labs control the pace and scale of the work. A boycott can register dissent; it cannot by itself give mathematicians the time, access or authority to shape what gets proved next.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X