Skip to main content
Home » Artificial Intelligence » News » OpenAI Claims an 88-Hour Maths Breakthrough, But the Controversy May Outlast the Proof

OpenAI Claims an 88-Hour Maths Breakthrough, But the Controversy May Outlast the Proof

6 min read
OpenAI Claims an 88-Hour Maths Breakthrough, But the Controversy May Outlast the Proof

Stay connected with KayaToday, follow us on Instagram and Facebook for the latest news and reviews delivered straight to you.


Solving a problem that has stumped mathematicians for nine decades in under four days sounds like the kind of headline that belongs in science fiction. OpenAI is insisting it belongs in the scientific record instead, and the claim is already generating as much friction as it is excitement.

On Tuesday, the ChatGPT-maker announced that it had deployed roughly 10,000 AI agents, meaning software bots that operate with a degree of autonomy to complete assigned tasks, on one of mathematics’ most notorious unsolved challenges: the Navier-Stokes existence and smoothness problem. Starting on 1 September, those bots worked for approximately 88 hours before producing what OpenAI is calling a solution. The company described the result as a “milestone” and evidence that its AI systems are advancing rapidly. Independent verification, however, has not yet arrived.

What Navier-Stokes Actually Is, and Why It Has Resisted Solution

The Navier-Stokes equations describe how fluids move, covering everything from air flowing over an aircraft wing to water churning through a river. The mathematics underpinning these equations has been understood in broad terms for well over a century, but a specific and deeply important question has remained open for roughly 90 years: do smooth, well-behaved solutions always exist for these equations in three dimensions, or can the mathematics break down into chaos under certain conditions?

This is not an abstract puzzle. Turbulence, one of the least understood phenomena in classical physics, sits at the heart of the problem. Engineers, meteorologists, and physicists all work around gaps in knowledge that a full proof would help close. The Clay Mathematics Institute in the United States recognised the problem’s significance by listing it among its Millennium Prize Problems, a set of seven mathematical challenges each carrying a prize of one million US dollars for the first verified solution.

OpenAI’s announcement is careful on this point. The company said its AI bots resolved two of the four statements required by the Millennium Prize proof. It explicitly stated that it does not intend to claim the prize for this result. The Clay Mathematics Institute has not publicly accepted or verified the work.

The Scale of the Computation, and What It Cost

The raw numbers behind the effort are striking on their own terms. OpenAI said the 10,000 AI agents exchanged nearly 3 million messages and generated 130 billion output tokens, which refers to the individual units of text and code that a language model produces in its responses, while working on Navier-Stokes alone. Based on OpenAI’s own published pricing for output from its most advanced models, that level of computation would have cost approximately 10 million US dollars.

The model used is not one available to the public. OpenAI said it began training this new internal model at the end of August and quickly found it to be unusually capable at mathematics. Researchers described it as “significantly more capable” than the company’s most recent public release. When rumours circulated on 1 September that two Millennium Prize problems had been resolved elsewhere, OpenAI decided to direct its new model and the army of AI agents toward the remaining unsolved problems. Four days later, it had results on Navier-Stokes.

A Parallel Effort and a Pointed Accusation

The announcement did not land cleanly. Within hours of OpenAI publishing its Navier-Stokes work on Tuesday, Tristan Buckmaster, a mathematics professor at New York University, made a public statement of his own. He said that he and Levent Alpöge, a mathematician employed by Anthropic, which is a direct competitor to OpenAI, had also been working toward a solution for the same problem. Their work had made use of OpenAI’s own Codex tool.

Buckmaster’s central allegation is one of timing and information flow. He stated that on 3 September, two days into OpenAI’s 88-hour sprint, he learned that “information about our progress had been passed to OpenAI.” He argued that OpenAI did not begin working on Navier-Stokes until after details of his and Alpöge’s work had reached the company. To support his account, he shared text from email exchanges with OpenAI about the matter. He was explicit about his motivation for going public, saying he felt compelled to share “what I was told, when, and what was proposed to me, because the alternative is to let a sequence of announcements say something I know to be false.”

OpenAI’s response was measured but firm. The company congratulated what it called the “remarkable” concurrent work of Buckmaster and Alpöge. It denied accessing any of their work through any channel before they released it publicly, and it denied using user data in its Navier-Stokes research. It did, however, leave a narrow opening, acknowledging that “while unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” It also noted that the two proofs differ significantly, including in the precise results they establish.

Why the Controversy Matters as Much as the Claim

The Navier-Stokes announcement sits at the intersection of two questions that will define how AI research develops over the next several years. The first is whether large-scale AI systems can genuinely advance the frontier of mathematical knowledge, not just assist human researchers but independently produce novel proofs that hold up to expert scrutiny. OpenAI is betting the answer is yes, and if independent mathematicians verify its Navier-Stokes work, that bet will look prescient.

The second question is about process and trust. Buckmaster’s account, if accurate, raises uncomfortable issues about how a company with access to vast user data and an internal model more powerful than anything publicly available should conduct research that overlaps with work being done by outside academics using its own tools. OpenAI’s denial is categorical, but the company’s own caveat about de-identified training data leaves enough ambiguity that the dispute is unlikely to resolve quietly.

For the broader AI field, both questions carry weight. A verified mathematical breakthrough of this magnitude would accelerate investment and confidence in AI-driven scientific discovery. A credible allegation of improper information use, even one that falls short of proof, would deepen concerns about the structural advantages that frontier AI labs hold over independent researchers who depend on those same labs’ products to do their work. The mathematics may eventually be settled. The ethics question will take longer.

Read More: Enterprises Are Burning Millions on AI Tools They Cannot Prove Are Working

Faraz Khan is a freelance journalist and lecturer with a Master’s in Political Science, offering expert analysis on international affairs through his columns and blog. His insightful content provides valuable perspectives to a global audience.
341 articles
More from Faraz Khan →
We follow strict editorial standards to ensure accuracy and transparency.