Research Strongly positive 9

OpenAI Claims 10,000 AI Agents Cracked Navier-Stokes in 88 Hours

OpenAI says an internal multi-agent system more powerful than GPT-6 Astra solved the Clay Millennium Navier-Stokes problem in 88 hours. But rival researchers question whether the company benefited from their unpublished work stored in OpenAI Codex.

· 4 min read · Verified by 2 sources ·

Beat this week

Last 7 days · Research

10 stories
6.5 avg impact
30% positive
40% negative
vs prior 7 days +8 +8 stories vs prior 7 days

Impact 6.5/10 (+0.5 vs prior). Counts are stories in our record, not a market forecast.

Open the change report

Coverage balance Negative coverage leads. Negative coverage exceeds positive coverage by 10 percentage points.

  • 30% positive
  • 30% neutral
  • 40% negative

This story sits in Research — the counts compare this beat's last 7 days with the previous 7 in our verified record, not a market forecast.

Figures are computed live from our source-verified story record (as of ) The volume change compares this window with the prior 7 days in the same record. — see our methodology for how impact and sentiment are derived.

AI briefing

Key takeaways

9 impact
Strongly positivesentiment
2sources
4min read
  1. OpenAI says an internal multi-agent system more powerful than GPT-6 Astra solved the Clay Millennium Navier-Stokes problem in 88 hours.
  2. But rival researchers question whether the company benefited from their unpublished work stored in OpenAI Codex.
Drawn from
  • theguardian.com
  • aol.co.uk

In this briefing

Mentioned

Key Intelligence

Key Facts

  1. 1OpenAI claims to have cracked the Navier-Stokes problem, one of seven Clay Mathematics Institute Millennium Prize Problems, after spending millions of dollars.
  2. 2Roughly 10,000 autonomous AI agents worked on the problem simultaneously and reached the claimed solution in 88 hours.
  3. 3OpenAI said the internal system used for the effort was more powerful than its latest GPT-6 Astra model.
  4. 4Tristan Buckmaster, an NYU professor, said OpenAI accelerated its work after hearing he and an Anthropic researcher were poised to announce their own breakthrough.
  5. 5Buckmaster said the pair's work in progress was stored in OpenAI's Codex model, potentially making it visible to OpenAI staff.
  6. 6OpenAI researcher Sebastien Bubeck denied using the rival work, but OpenAI could not rule out that data from the pair's use of its products "helped improve our models."

I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything.

Tristan Buckmaster Professor of Mathematics, New York University

In a document posted on his website

Analysis

For AI researchers, the headline isn't just a math proof—it's a trial run of massively parallel agentic problem-solving at 10,000-agent scale. The dispute over Codex-stored research also tests how training and product data boundaries hold up when frontier labs compete head-to-head on open scientific problems.

On 8 September 2026, OpenAI announced that it had cracked the Navier-Stokes existence and smoothness problem, one of the seven Clay Mathematics Institute Millennium Prize Problems, with an internal system that coordinated roughly 10,000 autonomous AI agents to reach a solution in 88 hours. The company said the effort cost millions of dollars and used a system more powerful than its latest GPT-6 Astra model. If independently verified, the result would be one of the most consequential AI-for-science breakthroughs in the field's short history. Within hours, however, the announcement was overshadowed by a dispute over whether OpenAI had benefited from unpublished work by a rival team at Anthropic.

The controversy centers on Tristan Buckmaster, a mathematics professor at New York University, who said OpenAI accelerated its work after hearing that he and an Anthropic researcher were poised to announce their own breakthrough on the problem.

The technical claim itself is striking. OpenAI says the 10,000 agents worked simultaneously, suggesting a shift from single-model inference toward massively parallel, agentic problem-solving. The 88-hour window is presented as evidence of the speed advantage of frontier systems over decades of human mathematical effort, and the disclosure that an internal system exceeded GPT-6 Astra implies OpenAI already has more capable systems in deployment than its publicly named model. But the full architecture, proof structure, and verification details remain unpublished, so the claim currently rests on OpenAI's account rather than peer-reviewed mathematics.

The controversy centers on Tristan Buckmaster, a mathematics professor at New York University, who said OpenAI accelerated its work after hearing that he and an Anthropic researcher were poised to announce their own breakthrough on the problem. Buckmaster noted that the pair's work in progress was stored in OpenAI's Codex model, which is used for writing computer programs, potentially making the material visible to OpenAI staff. His public statement was unusually careful: he said he does not know what OpenAI's model did, does not know whether his data was used, and was not accusing anyone of anything. At a press briefing the same day, OpenAI researcher Sebastien Bubeck denied that the company had used the pair's work or accessed material shared with OpenAI's servers. Yet OpenAI's own announcement added a caveat that could not rule out that data from the pair's use of OpenAI products "helped improve our models."

For the AI and mathematics communities, the episode raises immediate questions about provenance and verification. A correct solution to the Navier-Stokes problem would be a landmark event, but Millennium Prize problems require rigorous scrutiny before any claim is accepted. The Clay Mathematics Institute has not yet indicated whether OpenAI has submitted a formal proof. If the claim survives independent review, it would validate the use of agentic AI systems for open mathematical research. If it does not, or if the provenance dispute deepens, the episode could become a cautionary tale about frontier-lab competition, data separation, and the limits of announcement-driven science.

What to Watch

The dispute also highlights a structural risk in AI research platforms. Codex is a product used by outside researchers, and Buckmaster's concern is not that OpenAI stole code in the traditional sense, but that user work stored in a model or service could have informed a rival research program. That raises governance questions about how frontier AI companies separate product data from internal training and research, what audit trails exist, and whether competing researchers can safely use a rival's tools. OpenAI's statement that product data may have "helped improve our models" is likely to intensify scrutiny even if it falls short of admitting direct use of Buckmaster's work.

Looking ahead, the next phase will be defined by independent replication and release of the proof. If OpenAI publishes a complete, verifiable solution, the 88-hour, 10,000-agent result could accelerate investment in AI-for-science and multi-agent research systems. If the company withholds details or the proof cannot be verified, the controversy may overshadow the scientific claim. Either way, the incident provides a high-stakes test of whether frontier AI systems can be trusted to solve problems that humans could not, and whether the systems that attempt those problems operate under sufficiently clear data-governance rules.

Timeline

Timeline

  1. Clay Mathematics Institute publishes Millennium Prize Problems

  2. OpenAI claims Navier-Stokes breakthrough

  3. OpenAI press briefing denies wrongdoing

Source cluster

Primary reporting

2articles

Cite This Page

"OpenAI Claims 10,000 AI Agents Cracked Navier-Stokes in 88 Hours." AI Intelligence Brief, September 8, 2026. https://getaibrief.com/story/openai-10000-ai-agents-navier-stokes-88-hours

How we covered this story

Every story in our AI coverage is assembled from multiple primary sources, cross-referenced for factual consistency, and scored along three independent dimensions: sentiment, operational impact, and source-cluster confidence. Single-source rumors and unverifiable claims do not pass our editorial gate. When a story shows "Verified by N sources" with N≥2, the development is independently corroborated; when N=1, we mark it explicitly so readers can weigh the signal accordingly.

Impact scoring uses a 1-10 scale weighted toward regulatory, financial, and operational consequence rather than coverage volume. A topic that runs in every outlet but moves no real decisions ranks lower than a niche regulatory filing that reshapes how operators in the AI space have to behave. Read our full methodology for the scoring rubric, our glossary for term definitions, and our trends index for the longitudinal view across the beat.

Sources are only linked to a story once they clear our classification pipeline at a minimum 35 percent relevance threshold. According to that methodology, reviewed July 2026, this follows multi-source corroboration standards recommended by journalism research bodies such as the Reuters Institute for the Study of Journalism.

See something wrong in this story — a wrong fact, a broken source link, a misattributed entity? Report a data issue.