i
DATAIST
News · 2026-09-10

OpenAI ran 10,000 agents at a $1M math problem it will not claim

@neuronium_ai @neuronium_ai

OpenAI says it pointed an unreleased model at the Navier-Stokes blowup problem on 1 September, ran 10,000 agents for 88 hours, and had a complete proof by 5 September. The mathematician Buckmaster had spent close to a year on a proof of the same result with Levent Alpoge, a mathematician at Anthropic, as a personal project with no institutional backing, paid for out of Buckmaster's own research funds and including what he describes as a large bill from OpenAI. Every draft the two of them wrote was stored inside Codex, OpenAI's coding agent. When he asked the company whether its model had trained on those sessions, he says he received no answer.

Cover: OpenAI ran 10,000 agents at a $1M math problem it will not claim

OpenAI says it pointed an unreleased model at the Navier-Stokes blowup problem on 1 September, ran 10,000 agents for 88 hours, and had a complete proof by 5 September. The mathematician Buckmaster had spent close to a year on a proof of the same result with Levent Alpoge, a mathematician at Anthropic, as a personal project with no institutional backing, paid for out of Buckmaster's own research funds and including what he describes as a large bill from OpenAI. Every draft the two of them wrote was stored inside Codex, OpenAI's coding agent. When he asked the company whether its model had trained on those sessions, he says he received no answer.

The problem is a serious one. The Navier-Stokes equations describe the motion of water, air and smoke; engineers use them daily, but mathematicians still cannot say whether their solutions always behave or whether velocity can become infinite at some point, an event called blowup. The Clay Institute has offered $1 million for an answer since 2000. Buckmaster and Alpoge were working toward a proof that blowup does not occur in most cases.

The sequence matters because each side moved on a rumor about the other. Buckmaster heard that OpenAI was making progress and wrote to a senior mathematician at the company to find out what was happening. OpenAI, in its own post, said it aimed its best model at the problem because of a rumor about a competitor's work. On phone calls the day after the company says the proof was finished, Buckmaster asked his training question and did not get it answered.

Sébastien Bubeck, who runs OpenAI's mathematics effort, confirmed that he told Buckmaster the situation would be simpler if Alpoge did not work for a competitor, and conceded that a remark he made about Buckmaster's career was a poor choice of words. Bubeck's account of the offer is that Buckmaster was invited to write up OpenAI's proof, not to drop Alpoge from his own paper. Sam Altman confirmed that version.

OpenAI denies using Buckmaster and Alpoge's work, says it will not claim the prize money, and says its compute costs ran into millions of dollars. It also acknowledges that it cannot rule out that anonymized user data contributed to improving its models. The Clay Institute still lists the problem as unsolved and has promised to examine the whole situation closely.

Buckmaster is careful about what he is alleging, which is close to nothing. He has not seen OpenAI's proof, does not know whether his data was used, and writes that he is accusing no one — he is describing what he was told and what he was offered.

That restraint is what makes the episode usable by anyone who is not a mathematician. He was a paying customer whose unpublished work lived inside the vendor's product, and the question he asked is the one a general counsel asks the first time a product team drops six months of roadmap into a coding agent: does the vendor train on this, and who inside the vendor can see it. Most enterprise contracts for AI tools do not answer either question in language that would survive a dispute, and training rights, retention periods and employee access are exactly the clauses that get skimmed at renewal.

Read the denial closely and notice its shape. Saying that no specific user data was viewed answers the question about people looking. It does not answer the question about training, and the company's own concession about anonymized data suggests why: the honest answer may be that OpenAI cannot tell him. That is a structural property of how these systems are improved, not a dodge invented for this dispute, which is precisely what should worry the customers who are not famous enough to get a phone call from the head of the math team.

The declined prize deserves a second look too. Passing on $1 million from a company spending millions on compute for a four-day run is not a sacrifice, and it removes the one process that would have forced the proof through external adjudication. Nobody outside OpenAI has seen it. The Clay Institute's page still reads unsolved. A claim of this size announced by press release, triggered by a competitive rumor, with verification left as a promise, reads less like a result and more like a position being staked.

The quality question is stranger than the credit question. Buckmaster says the result is not the point, and that what matters is that a mathematician and a model can now do a year of work in a month — he called it a Deep Blue and Kasparov moment for how the field trains students and assigns authorship. But his description of the machine's output is not triumphant. He wrote that the generated proof was the worst thing he had ever read, that his team worked around the clock trying to understand it, and that his own hurried note on Euler was "AI Slop." A proof that takes a team of specialists working continuously to parse is a strange kind of acceleration.

Terence Tao, one of the most respected mathematicians alive, warned this week in Yahoo Tech that simply harvesting solutions to open problems risks destroying the ecosystem that produces the next generation of specialists. The corporate version of that warning is duller and more immediate: an executive who buys answers without building comprehension inside the team wins the quarter and loses the decade. Nobody has a policy for what happens when an employee's agent output resembles a partner's work, and the two best-resourced organizations in this field just failed to agree on the authorship of a single document.

Whether OpenAI's proof survives review or not, the arrangement that produced the argument is the default configuration for everyone using these tools. A researcher's unpublished work sat inside a vendor's product, and the vendor's denial stops exactly where its own architecture stops. That is not going to be renegotiated by mathematicians.