On August 27 I published They Said the AI Solved Ten Unsolved Problems. It Had Looked Them Up. That piece also carried a warning aimed at its own argument:
"It only searched the literature" was correct about that specific incident. As a general claim about AI and mathematics, it is now false, and anyone still using it is goalpost-shifting.
That warning got its hardest test on September 8, when OpenAI announced what it calls a solution to the Navier–Stokes problem, one of the $1 million Millennium Prize Problems. This claim did not collapse. It arrived with a 166-page manuscript and a formal proof in Lean, a programming language used to check whether proofs are correct, and the code compiled. Based on that, Javier Gómez-Serrano of Brown University told NPR, "the community seems to have the consensus that it is correct." The Clay Mathematics Institute, which runs the prize, says the problem has "apparently been settled." By the test my August piece proposed, "Judge the claim by what shipped with it," this one passes (with two asterisks: which version of the problem it settles, and who, by name, has checked it).
So the argument moves from whether the proof is true to what a proof is for. On that, NPR's summary of what mathematicians say is blunt: "so far, humanity has learned very little" from this one.
So what did the machine prove, exactly?
The Navier–Stokes equations describe how fluids such as water and air flow. The prize question accepts a proof of any one of four statements. Two, (A) and (B), require the outside force on the fluid to be "identically zero." The other two, (C) and (D), concern the equations breaking down, and they allow a smooth push from outside.
OpenAI's construction takes a fluid at rest, pushes it with a smooth outside force, and ends with speeds "growing without bound within a finite amount of time." Its manuscript says this "establishes alternative (C)," and the company says it covers (D) too. Read literally, a correct proof of (C) settles the problem as the prize posed it.
It is not, however, the question most mathematicians meant — whether a fluid left entirely alone can break down. Luis Silvestre of the University of Chicago told Scientific American that "The Clay problem is settled, but the main problem for the Navier-Stokes equations is not." In the same article Diego Córdoba, a Madrid mathematician whose earlier work this proof builds on (so not a neutral voice), pushed back: "All fluids we know of are under some kind of external force."
How was it made? With an internal model OpenAI describes as significantly more capable than its public GPT-6 Astra, run as roughly ten thousand agents at once for about 88 hours, plus 17 more hours for Astra's Lean version. Using OpenAI's prices for its most advanced public model, NPR put the computing for this run at "somewhere around $6-$10 million"; other estimates conflict.
And the machine check? It has a last step that belongs to people. As Quanta Magazine put it, humans must still "guarantee that the statement being shown to be true in Lean is logically equivalent to what mathematicians set out to prove." OpenAI's file takes its reference statements from an independent project, Google DeepMind's Formal Conjectures — a real safeguard. Yet the same file lists its review status as "self-assessed", and in Science News on September 9, the mathematical physicist Gregory Eyink of Johns Hopkins said he doesn't think anyone has completely verified the proof yet, "certainly not on the human side." As of September 30, Clay files the problem under "Active problems," apart from its solved and unsolved lists.
That's the certitude side. On understanding, the early reviews are rough. "So far it's been very difficult to really extract any human understanding from this new AI proof," Oxford's James Maynard, a Fields Medalist, told NPR. Gómez-Serrano was blunter: "The paper is not written for humans." He believes it could help advance the field after "some serious re-writing," but "as of today, the paper doesn't teach us much."
Whose idea was it, anyway?
Quanta reported that OpenAI and a rival pair of mathematicians both leaned heavily on a strategy by Córdoba and Luis Martínez-Zoroa of CUNEF University that "radically departed from the methods most mathematicians were using," and judged that "the intellectual debt to Córdoba and Martínez-Zoroa seems clear."
The rival pair is where the fight starts. Tristan Buckmaster, an NYU mathematician, had spent almost a year on the problem with Levent Alpöge, using models from both OpenAI and Anthropic. Alpöge works at Anthropic, which also makes Claude, the AI model that helped draft this piece; TechCrunch reports that he "was not conducting this research on the company's behalf." (Ties run every direction in a field this small, and none of them is an accusation: Buckmaster and Gómez-Serrano co-authored a 2025 paper on fluid singularities with Google DeepMind researchers.) Anthropic uses famous theorems as showcases, too — four days before OpenAI's news, it announced a computer-checked proof of Fermat's Last Theorem, noting that "what's novel here is the verification."
OpenAI says its effort began on September 1 after hearing "a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster," and it recognizes "the priority of their work on forced Euler," a closely related problem, and congratulates the pair.
Buckmaster set out what he says he was told in a statement that calls the pair's own Euler write-up "AI slop" and says the reasons "involve our being pressured by outside factors." By his account, OpenAI's second proposal was that he alone write up its Navier–Stokes result, and OpenAI's Sébastien Bubeck "twice asserted that he wanted Levent removed from authorship." He says that when he asked whether the model had been trained on, or had access to, their Codex sessions, which held "all our drafts," he was told it did not look up user data, and that when he asked again about training, he "did not get an answer." He also lists what he is not claiming, including: "I do not know whether our data was used. I am not accusing anyone of anything."
OpenAI's page answers: "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem." Bubeck called the claims against him "false and inflammatory allegations," then expanded: he "never ever asked for Levent to be removed from authorship of his own work"; he said "it would be simpler if Levent was not an Anthropic employee" while discussing Buckmaster leading a rewrite of OpenAI's proof; and "it was admitted that internal Anthropic models had been used in their proof of Euler blowup."
On training, the company's answer changed. On September 8 it wrote that it "cannot rule out" that "de-identified data derived from their usage of our products helped improve our models," while calling that "unlikely." On September 10 its own investigation's finding replaced that line: Buckmaster's Codex prompts from the two months before the announcement "could not have influenced the system in any way, including through training."
The manuscript's title page lists a single author: OPENAI. The most generous line came from Madrid. "It would have been nice to do this ourselves, but I'm very happy for him," Martínez-Zoroa told Quanta, of Buckmaster.
Isn't "doesn't teach us much" just the next goalpost?
Fair question, given what my August piece said. The strongest case for yes comes from inside mathematics. Timothy Gowers, who chose not to sign the Fields Medalists' letter "despite agreeing with much of what it said," notes that "what AI is producing is not just true/false statements," and that the old worry about "utterly opaque proofs" has "not turned out to be the case," even if the write-ups "often leave plenty to be desired." (He discloses contacts at OpenAI, early access to some of its models and free access to its Pro models, and says he has never been paid by OpenAI.)
So why isn't this goalpost-shifting? Because this goalpost was there before the game started. The prize's own page says so: "Why ask for a proof? Because a proof gives not only certitude, but also understanding." Terence Tao of UCLA, writing about Buckmaster and Alpöge's work the day before OpenAI's announcement, called solving such problems "only a proxy goal" for "developing mathematical understanding and insight."
Silvia De Toffoli and Eamon Duede, writing on the blog of Princeton's Center for Information Technology Policy, name the distinction: "two notions of proof: a logical notion and an intelligible notion." They grant that "if OpenAI's announcement is correct, this is an extraordinary achievement," and that the Lean file "secures certainty." But "At this moment, it is an answer, not a solution." And, pointedly: "This is not moving the goalposts but recognizing that any specific goalpost is inadequate."
Mathematics has lived with proofs it couldn't check before
Has the field been here before? Twice, famously, and both times the other way around.
In 1976, Kenneth Appel and Wolfgang Haken completed a proof of the Four Color Theorem (four colors are enough to shade any map so that no two adjacent regions match) by reducing it "to a finite (albeit large) number of cases" and handing those to a computer. The University of Illinois calls it "the first computer assisted proof of a major theorem" and notes that the "considerable initial criticism" eventually abated. The Urbana Post Office issued a postmark reading "four colors suffice." (Yes, a postmark. For a theorem.)
Did understanding follow? Slowly, and not all the way. In 1998, Robin Thomas wrote in the Notices of the American Mathematical Society that Appel and Haken's work was "a major breakthrough in mathematics," but that "there remains some skepticism regarding the validity of their proof" and that the theorem, "even today," was "not yet fully understood." He and three colleagues had "tried to verify the Appel-Haken proof, but soon gave up" and worked out their own. Georges Gonthier, whose team finished a complete formal proof in 2005, later wrote that formal proof is not only a way to rule out mistakes "but also a tool that shows us and compels us to understand why a proof works."
The Kepler conjecture, about packing equal-sized balls as densely as possible, went the same way. Thomas Hales's computer-assisted proof (announced with Sam Ferguson) reached the Annals of Mathematics on September 4, 1998, and was accepted on August 16, 2005. An editor's letter, as Hales later quoted it, said the referees "have not been able to certify the correctness of the proof"; Hales, hardly a neutral narrator, has since written that "the referees became exhausted and quit." So he started Flyspeck, a collaboration that completed a formal proof in 2014, published in 2017. Its paper asks the question that now matters most for Navier–Stokes: "did the right theorem get formalized?"
Here the direction flips (my reading, not anyone's finding). In 1976 and 1998, humans had an argument they could follow and a machine-sized part they couldn't check; De Toffoli and Duede note that Flyspeck was partly meant to confirm an "intelligible (but hard to check) proof." This time the machine check came first, in 17 hours, and the intelligible part has barely begun: a readable version of one part appeared September 28. Gonthier's team got understanding as a by-product of formalizing a proof themselves; here a machine did that step, so whatever it taught passed through no person. The Four Color story supports both readings: understanding that catches up, and understanding that lags for decades.
Now picture the next hundred answers
What happens when this stops being one proof and becomes a pipeline? OpenAI says the same model "has now resolved more than 100 long-standing open problems across most areas of mathematics," though it had published no list by September 30. Nine unpaid mathematicians formed an independent advisory group after OpenAI approached some of them; its current task is advising OpenAI "on how to coordinate the release of a large number of significant results in mathematics that they report have been produced by their internal model." On September 29 the group urged AI labs not to use releases as "marketing vehicles" and, for results nobody understands yet, to "take responsibility for ensuring that human understanding will follow."
Gowers, one of the nine, sketches the scenario he considers more likely: answers arriving from public models at a rate that "far exceeds the rate at which the mathematical community can absorb them," most of them obtained "with zero effort from human mathematicians," after prompts like "Thank you — please continue."
Now put yourself in it. You are a doctoral student in 2028, two years into a problem chosen to teach you the field. One morning it arrives answered — in a manuscript bylined to a company, checked by a machine, understood by no one on your committee. The "primary risk," as Gowers sees it, is that a lot of people who would have become "custodians of the mathematical tradition will no longer wish to do so."
Then picture the first Clay committee to face a file like that. Clay's rules describe a winner as "one person" or "multiple solvers of a Problem or their heirs" (yes, heirs), and say the institute "will pay special attention" to whether a solution "depends crucially on insights published prior" to it. Who is the solver when ten thousand agents did the work and two Madrid mathematicians developed the strategy? Who, exactly, are an agent's heirs? I won't guess how Clay would answer.
And picture what it does to how mathematicians talk. Tao has warned that "The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community," which "would reverse centuries of traditions of open science." In that world, the most important switch in a mathematician's software decides whether unpublished drafts become training data.
Who is worried, and about what
On the market-oriented side, AEI's James Pethokoukis, looking at a study in which many scientists say AI pushes them toward "safer, more incremental projects," calls that "Not great if your ultimate goal is pushing forward the science frontier," then adds: "Still, early days for this powerful general-purpose technology."
From the center-left, Michael J. Ahn, writing for Brookings, warns that once AI can produce "a working proof," the artifact "stops being what a university can meaningfully grade on its own," and universities "should not try to out-factory the machine." (Brookings discloses Google as a donor.)
Institutional mathematics has answered, too. A September 11 declaration, first signed by twenty-five Fields Medalists and naming no company, says the goals of AI companies and mathematicians are "severely misaligned" while granting that "AI offers the potential of enhancing and accelerating genuine mathematical study and understanding"; as of September 30 it lists 8,118 endorsers.
On the critical end, Columbia's Michael Harris, in comments to Science he reposted on his newsletter, worries the highly publicized achievement will be "extremely damaging to mathematics," saying it "convinces young people that their passion for mathematics has no future."
What does this mean for you?
Before you share a headline that says AI solved something, ask which version it solved. This prize accepts four statements; OpenAI's proof is of the kind with an outside push, not the force-free version most mathematicians meant.
Read a Lean check as what it is: a machine checking a formal statement. Then ask who confirmed it matches the real question, and who outside the company reviewed it. For now, the company's own file says "self-assessed."
Ignore any headline that says someone won the $1 million. Clay's rules require publication in a qualifying outlet, two years and general acceptance before the institute will even consider a solution, and OpenAI says it does not intend to claim the prize.
If you do original, unpublished work in a consumer AI tool, check its training setting today. In ChatGPT, turn off "Improve the model for everyone" under Settings → Data Controls, and check Codex's separate setting. In Claude, it is "Help Improve our AI models"; in Gemini, "Keep Activity." All three switches work going forward, so flip yours before your drafts go in. None of this implies anyone's data was misused here: OpenAI says its investigation ruled that out for two months of Buckmaster's Codex prompts, and he says he does not know.
Keep the "apparently" when you pass it on. NPR wrote that OpenAI "apparently solved" the problem; Quanta's article says "If the result holds up." My August test still applies: "Ask who checked, by name."
If you do research, borrow Leiden's language. The Leiden Declaration, which the International Mathematical Union endorsed, says material "should not be used as training data without consent," and asks policymakers to consult experts "rather than relying on press releases or popular reporting of mathematical results." That last line is worth sending to your representative.
The lesson, as I see it
This claim shipped the goods: a manuscript, a public formal proof, and a result few mathematicians believe is wrong. That deserves to be said without a "but."
Then comes the "and." By the prize's own account, a proof owes you understanding as well as certitude, and here the understanding has barely begun to arrive. History says that can take decades, and that it arrives through people.
Clay says it hopes for "waves of new human understanding unleashed as the innovations behind this work are analysed and interrogated." So do I. My vote? Celebrate the certitude, fund the understanding (the advisory group says labs should help pay), and ask of every lab what Buckmaster asked of this one: if an OpenAI model did close the gap, "that is a remarkable thing and it should be said loudly, by them, with the history intact." That history starts with two names from Madrid.
The machine passed the exam. Someone still has to teach the class.





