OpenAI said on Tuesday that a swarm of its agents had produced a proof for the Navier-Stokes equations, a 90-year-old fluid-dynamics problem that carries a $1 million Clay Mathematics Institute prize. The coordination layer alone exchanged 2.7 million messages across roughly 10,000 concurrent agents, burning through 130 billion tokens in 88 hours before a further 17 hours of verification. At consumer pricing, the output tokens imply a $6.5 million bill; including the larger input volume pushes the estimate to $10 million to $40 million. The expense makes the result a live test of whether frontier reasoning can justify its own compute budget.

The token accounting

LisanBench, an independent benchmark evaluator, derived the cost range from OpenAI’s published token counts and its average consumer rates. The firm did not disclose whether the run used discounted internal pricing or a dedicated cluster, so the $40 million ceiling is a retail proxy, not an audited invoice. Sam Altman responded to the estimate on X by calling AI a bubble, a line that reads less like a joke when the company’s own chief executive is the one pricing the inference.

The verification gap

OpenAI has not released the proof, the agent transcripts, or a formal write-up for peer review. The 17-hour verification window was internal. Until a qualified mathematician or the Clay Institute examines the artifact, the claim sits in the same category as any unreproduced benchmark: plausible, expensive, and unverified. The prize rules require publication in a refereed journal and two years of general acceptance; neither condition has been met.

The training-data dispute

Mathematicians Tristan Buckmaster and Levent Alpöge questioned whether OpenAI’s models had absorbed their related research through product-usage data. OpenAI denied direct access to their work but acknowledged it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.” The phrasing leaves the door open for any user-generated content to become training material unless explicitly opted out, a policy that turns every chat session into a potential data contribution.

What to watch

The Clay Institute has not commented. If the proof withstands scrutiny, the $1 million prize will cover a fraction of the compute bill. If it does not, the episode becomes a costly demonstration that token volume is not a substitute for mathematical rigor. Either way, the episode establishes a new unit of account for AI research: the 130-billion-token experiment.