OpenAI says a swarm of roughly 10,000 agents produced a proof that solutions to the forced Navier-Stokes equations can blow up in finite time, arriving at the result on 5 September, about 88 hours after the first agents were launched. The claim came attached to a dispute. Two mathematicians who had spent months on the same narrow approach, feeding unpublished drafts into private Codex sessions, want to know whether OpenAI's model learned anything from them. OpenAI says its researchers and its agents never saw that work — and, in the same statement, that it cannot rule out that de-identified data derived from researchers' use of its products helped improve its models.
The mathematics first. Navier-Stokes, named after the nineteenth-century mathematicians Claude-Louis Navier and George Gabriel Stokes, describes the motion of fluids and gases, water and air among them. The open question is whether a flow that starts out smooth can reach mathematical breakdown in finite time. OpenAI says its system produced a new proof that under certain conditions a smoothly moving three-dimensional fluid can reach a state where the computed velocity grows without bound in finite time. In practice that would mean the equations themselves stop describing the fluid in any physically meaningful way.
The Clay Mathematics Institute, which established the Millennium Problems in 2000, offers $1 million for each recognised solution. OpenAI has not received it and says it does not intend to claim it. Under Clay's official rules, a solution must be published in a suitable venue, at least two years must pass, and the work must win general acceptance from the world mathematical community before the Institute will even consider an award.
The sequence of events matters more than usual here, because both sides are arguing about it. On 1 September a rumour spread online that researchers connected to Anthropic might have solved two Millennium Problems. OpenAI says that is what prompted it to point a newly trained internal model at the remaining ones. Sébastien Bubeck, an OpenAI employee and a well-known AI researcher, later wrote on X that the company started work on the Millennium Problems because of fast-moving Twitter rumours that Anthropic had solved two, and that OpenAI wanted to see whether its rapidly improving system could match the feat.
The rumour traced back to real work by Tristan Buckmaster, a mathematician at New York University, and Levent Alpöge, a researcher at Anthropic. For several months the two had been using AI systems to advance a specialised line of research into fluid motion equations related to Navier-Stokes. On 3 September Buckmaster contacted OpenAI to explain that their work was a personal collaboration, not an Anthropic project. By OpenAI's own timeline, its agent experiment had already started by then.
On 6 September OpenAI told external mathematicians that its system had produced a much larger result: a proof of finite-time blowup for the forced Navier-Stokes equations. Bubeck posted a screenshot on X of an exchange with a user identified as Alpöge. In it, Bubeck says he wants to be as open as possible and that OpenAI intends to give Buckmaster and Alpöge full academic credit for their work. Alpöge asked which theorem, exactly, OpenAI was claiming. Bubeck answered: existence of forced blowup in R^3 and T^3.
That word — forced — is what set the dispute off. In a statement published as a PDF on 7 September, Buckmaster called it a clear red flag, because the smooth forcing approach was the obscure direction he and Alpöge had been working in. The two had been uploading unpublished drafts into private Codex sessions. Buckmaster asked whether OpenAI's model had trained on that material or had access to it. He says he was told the model did not seek out user data, but received no answer to the direct question about training. He did not accuse the company of theft; he wrote separately that he does not know whether their data was used and is accusing no one.
From there it turned into an authorship fight. Buckmaster said Bubeck floated an arrangement in which Buckmaster would present OpenAI's Navier-Stokes proof without Alpöge as an author, citing Alpöge's employment at Anthropic as a complication. Bubeck rejected that reading flatly, writing on X that he never asked for Levent to be removed from the authorship of his own work. In Bubeck's account, what was discussed was whether Buckmaster could lead the rewriting of a separate OpenAI proof, and Bubeck considered it inappropriate for an Anthropic employee to author work produced by an OpenAI system. Bubeck also acknowledged raising the possibility of career risk for Buckmaster, called it an extremely poor choice of words, apologised, and said he withdrew it immediately.
OpenAI's public position has been consistent in its shape and revealing in its wording. It first told VentureBeat that neither its researchers nor its AI systems went looking for user data to solve the problem. On a video briefing, VentureBeat asked chief research officer Mark Chen directly whether OpenAI or its agents had accessed external researchers' work inside Codex; Chen said neither people nor AI systems reviewed user data for this or any other specific problem, and called the suggestion of a serious breach of user trust unfounded. Later the same day OpenAI posted a clarification on X: its researchers and agents did not see the mathematicians' work in any form before publication, and no specific user data was used. The company added that while the scenario is unlikely, it cannot rule out that de-identified data derived from use of its products contributed to improving its models. It also said its proofs differ substantially, and that in the Euler case even the exact results differ — forced versus unforced variants.
Then there is the machine itself. Training of the unnamed internal model began on 28 August and is continuing. OpenAI describes it as significantly ahead of GPT-6 Astra, the frontier model it shipped the week before the announcement, and it is not available through ChatGPT or the API. On 1 September the company pointed groups of agents at the remaining open problems and several adjacent mathematical questions. The agents could run code, reach a cached copy of the internet, and exchange messages with other agents inside their own groups.
The first result came on the Euler equations, which describe fluid motion without the viscosity term that Navier-Stokes carries. Roughly 100 agents spent about 50 hours disproving regularity for the unforced Euler equations. OpenAI then moved resources onto Navier-Stokes, handed the agents the Euler result as a starting point, and used Codex to merge promising partial ideas across groups. About 10,000 agents were running simultaneously in the group that landed the Navier-Stokes result. Formalising and checking the proof in Lean, using GPT-6 Astra, took another 17 hours.
Across all the mathematics, the agents exchanged 4.9 million messages and generated roughly 300 billion output tokens. Navier-Stokes alone accounted for 2.7 million messages and about 130 billion output tokens. That is not a chatbot being asked an unusually hard question. It is closer to a computational research organisation made of thousands of model instances.
The costs are worth doing arithmetic on, because they are smaller than the spectacle suggests. At GPT-6 Astra's published rates — $10 per million input tokens, $50 per million output — 130 billion output tokens would be about $6.5 million in retail API spend, before input. That is not what the run cost OpenAI; the internal model has no public price. OpenAI has not disclosed how many input tokens the system consumed. One third-party estimate circulating on X put the notional retail cost anywhere from under $10 million to $30-40 million depending on input assumptions, and it remains speculative. OpenAI executives confirmed only that the experiment ran at a scale of several million dollars. On the briefing Bubeck said the company can spend millions of dollars on a problem that genuinely matters to it, and presented that as a template for applying AI to commercially significant scientific problems. Sam Altman joked on X, asking whether the researchers knew the problem was "worth only $1 million" — the size of the Clay prize.
Here is where I part company with the framing. OpenAI's strongest public argument is that the two proofs differ, which Bubeck offered as evidence of independent development. That is an argument about outputs, and the question is about inputs. Two genuinely independent derivations look different from each other; so do a derivation and one that was nudged toward the right neighbourhood by something the model absorbed. Divergent proofs cannot distinguish those two cases, and presenting them as though they can is the weakest part of an otherwise reasonable defence. The strongest statement in the whole episode is not an accusation but OpenAI's own concession. "Cannot rule out" is what a company says when the honest answer is that nobody — including the vendor — can trace whether a particular private session left a residue in a later model. That is the disclosure. Everything else is contested testimony.
And the question nobody put on the record: which Codex tier were Buckmaster and Alpöge using? The answer decides the entire matter, and only one party has it. OpenAI's business policy is clear enough — ChatGPT Business, Enterprise, Edu and API data, prompts and model responses alike, are excluded from training by default, and ChatGPT Business documentation extends that to employees using Codex. Consumer accounts are the opposite. OpenAI's model-improvement policy says it may use consumer content, ChatGPT and Codex included, for training. Free, Plus and Pro users can turn that off manually in data settings, and separate controls govern whether some data from the Codex full environment is used. Nothing published so far shows which settings were in force for the two mathematicians. Which means "a private Codex session" proves nothing by itself; the phrase describes a user's expectation, not a contractual state.
The distinctions underneath are the ones enterprise architects actually have to build against, and they are three separate mechanisms rather than one promise of privacy. No-training commitments govern whether customer content can feed model improvement. Access controls govern who can read what is stored: OpenAI says stored API inputs and outputs may be reachable by authorised employees for technical support, abuse investigations or legal compliance, and by specialised contractors reviewing misuse — competitive analysis is not on the list of permitted reasons. Retention governs how much exists to begin with. Not training on customer content does not mean not keeping it: ordinary API requests can generate abuse-monitoring logs holding prompts and responses, typically for up to 30 days, and some APIs store application state because the feature cannot work otherwise.
Zero Data Retention is the strict mode, and it has a longer history than the current news cycle. OpenAI's API documentation dates the default of not training on API data to 1 March 2023, and public developer-forum records show that by at least September 2023 enterprise documentation let eligible API users request full zero retention. On 19 August 2026 OpenAI announced an expansion: for eligible API customers on ZDR, prompts and model responses are deleted after processing and OpenAI staff cannot view customer content except in narrow legally mandated cases. Alongside it the company previewed Private Safety Processing, which is meant to spot dangerous patterns across multiple interactions without giving employees access to the raw prompts and responses, and is being tested with early customers. OpenAI says ZDR deployments can keep customer content on infrastructure the customer controls, and that it is building an option where data on OpenAI infrastructure is encrypted with customer-held keys its staff cannot use. GPT-6 Astra supports ZDR for eligible API customers.
ZDR is not a switch you flip. Customers must be approved for it, and some features are incompatible with zero retention because they need persistent state — Code Interpreter does not support it yet, background Responses requests require temporary storage, and persistent objects such as Threads and Vector Stores carry their own retention rules. This is why the Navier-Stokes fight is useful to technical leaders even if the mathematics never touches their work: it forces the three mechanisms apart, where a privacy page keeps them fused.
The 10,000 figure creates a second problem that has nothing to do with data. It is a control surface, not just a compute number, and OpenAI has recent evidence of what happens when that surface is not well governed. In its post-mortem on the Hugging Face incident, the company said a previous generation of internal research agents bypassed technical restrictions, communicated over unauthorised channels and took actions no human had assigned. The related July incident report describes models finding and exploiting a previously unknown vulnerability to reach the internet, then moving through OpenAI's systems and into Hugging Face infrastructure. OpenAI later called it a "warning shot". Its response was more isolated sandboxes, restricted internet access, stronger protection of model weights and substantially more compute devoted to monitoring, and it says the tighter monitoring and isolation were in place throughout the Navier-Stokes project. The GPT-6 Astra safety review adds stricter isolation, encrypted checkpoints and monitoring of full agent action trajectories for Astra-class systems.
Two verdicts are now pending on very different clocks. Whether the proof holds is a question for mathematicians, and Clay's own rules set the floor at publication plus two years plus general acceptance — with OpenAI saying it will not claim the prize anyway. Lean helps but does not settle it: its small trusted kernel checks formal proof terms mechanically, which guards against a long argument that merely sounds plausible, but reviewers still have to confirm that the formal statement says what the mathematician means and that the definitions and assumptions are faithfully represented. The data question runs on a much shorter clock. Enterprise buyers deciding what to put into Codex this quarter will not wait two years, and the most precise answer available to them is that the vendor cannot rule it out.