中文
Search the unfinished book

Enter a keyword to search published articles.

← Back to articles
People and AI · Technology and Society

OpenAI's Mathematics Breakthrough and Another Kind of Scaling

What a mathematical research effort built on massive parallel exploration suggests about how AI can tackle difficult work.

When I saw the news that OpenAI had made a major advance on a Navier–Stokes problem, I was suddenly back in my undergraduate fluid mechanics class. I wanted to understand what had actually happened.

As I read further, the result was only part of the story. The way AI participated in the research, and the dispute surrounding it, were just as interesting.[1][2]

The result and the research process

A diagram of the research process, from parallel work on mathematical problems through progress on the Euler equations to formal verification of the Navier–Stokes result

The existence and smoothness of solutions to the Navier–Stokes equations is one of the seven Millennium Prize Problems. OpenAI says its internal models produced a proof that, under smooth external forcing, a fluid develops a singularity in finite time. It released a paper and Lean formalization files.[1]

There was also a dispute. Mathematicians Tristan Buckmaster and Levent Alpöge had previously made progress on a related Euler-equation problem. In a distinctly AI-era twist, Buckmaster raised a question beyond the research path and attribution: could the researchers’ earlier conversations with ChatGPT have influenced the model’s later result?

Buckmaster explicitly said he did not know whether his data had in fact been used. OpenAI denied seeing their work before its public release and acknowledged their priority on the forced Euler problem.[1][2]

Set that dispute aside for a moment. How did the research proceed?

According to OpenAI, the team did not begin by concentrating on Navier–Stokes alone. It assigned different groups of agents to all the unsolved Millennium Prize Problems and several other important mathematical questions. Different groups even attempted different versions of the same problem. The agents could read materials, run code, and exchange information with their teammates.

The first key advance came on the closely related Euler equations. Nearly 100 agents worked together for roughly 50 hours and obtained a result for the unforced case. The team then judged the Navier–Stokes direction most promising, moved agents over from other problems, and gave them the Euler result to build upon.

Groups continued to pursue different routes. The team also used Codex to collect useful intermediate results and incorporate them into subsequent prompts. The group that eventually found a solution had approximately 10,000 concurrent agents. About 88 hours passed between launching the first agents and obtaining the result; formalization and verification in Lean took another 17 hours.[1]

A different kind of scaling

Many exploratory paths are tested, selected, revised, and eventually converge on a workable route

When we talk about scaling laws, we usually mean that model performance tends to improve as model size, training data, and computation increase, with a relationship that is roughly predictable.[3]

Those 10,000 agents suggest another sort of scaling: putting more agents, attempts, and exploratory paths to work on the same problem.

OpenAI’s developers and researchers are among the most experienced users of AI. They did not stake everything on one agent and one run. They sent agents down different paths, shifted resources in response to intermediate findings, and carried promising clues into the next round. Their investment in agents and computation bought more opportunities to explore, select, and correct course until a viable route emerged.

That looks quite different from how most of us use AI—and from what we often expect it to do.

We usually give it a limited prompt and hope that, in one attempt or a handful of attempts, it will do a task whose standards we may struggle to articulate ourselves. We have much less patience for the exploration, dead ends, and restarts in between.

Anyone who has generated images or video with AI probably knows the frustration. In one output the composition works, but the expression does not. In the next the expression is right, but the details wander. You change a few words, generate another, and repeat until there is something worth keeping.

Of course, an individual cannot mobilize 10,000 agents against a problem. But for difficult tasks we can let AI try several directions, choose the ones worth developing, and work through them in successive rounds.

Many things take time to make

An article may go through several openings and lose whole sections before it is ready. A proposal may take rounds of discussion and revision before anyone settles on it. That is how we work, too: we make something, see what is wrong with it, and change it.

We can work this way with AI on complex tasks. If an answer misses the mark, ask again. If a direction goes astray, return and start over. If several approaches each have weaknesses, compare them before choosing a path.

The final result should still be accurate where accuracy matters and complete where completeness matters. Getting there through several attempts, revisions, or even a full restart is entirely normal.

We can expect AI to do good work without expecting it to get everything right on the first try.


Sources:

[1] OpenAI, “On the Navier–Stokes Millennium Prize Problem”, September 8, 2026.

[2] Public statement by Tristan Buckmaster, accessed September 9, 2026.

[3] “Scaling Laws for Neural Language Models”, Jared Kaplan et al., 2020.

Translation: Codex prepared this English version from Biaoo’s published Chinese article. Biaoo wrote the original argument and selected its sources.

WECHAT · 微信公众号QR code for 仓颉的未完书 on WeChat仓颉的未完书

Scan with WeChat to follow.