I forwarded it.
That is the confession, and I want it on the table before I say anything clever. In October 2025 an OpenAI vice-president posted that GPT-5 had found solutions to "10 (!) previously unsolved Erdős problems and made progress on 11 others" — the Erdős problems being a famous catalogue of conjectures left behind by one of the twentieth century's most prolific mathematicians, the kind of thing careers get spent on. I read it. Something in my chest did the little lurch it does when a line in the sand moves. And I sent it to three people I respect with a message roughly equal to look at this.
Within days, the mathematician who maintains the database called the claim "a dramatic misrepresentation." The model had not solved anything. It had found existing published papers — real solutions, by real humans, sitting in the literature — that the database had not yet catalogued. The OpenAI vice-president deleted the post. Google DeepMind's chief executive called the episode "embarrassing." One technology news site summarized the whole affair as a breakthrough that never happened.
What stung was not the company's error. It was that I had been an amplifier. Three people got a thrilling and wrong thing from me, and I can tell you with complete confidence that I did not send them the correction with the same energy. Nobody does. Retractions travel economy class.
So I did the work I should have done first. The honest answer is more interesting than either "AI solved ten unsolved problems" or "AI can't do maths" — because since that deleted post, an AI system genuinely did disprove an eighty-year-old conjecture, and nine mathematicians put their names to checking it. The problem was never the machine. It is the gap between what gets claimed and what gets verified, and how many weeks apart those two things happen.
First, the boring order of events
The whole thing rested on a website. Not a journal, not an institution — a website one mathematician maintains as a labor of love. As of late July 2026 it lists 1,217 problems, of which 565 (46%) are recorded as solved. It is a wonderful public good, and it is also, by construction, one person's reading list.
And that matters enormously, because of what the word "open" means on that site. It does not mean humanity has not solved this. Its maintainer's own gloss, quoted afterwards in the coverage, is that "open" simply means he has not seen a paper that solves it. GPT-5 performed superbly at the task it was actually given; it just was not the task everyone thought. As one write-up put it, the model showed real capability — not as a mathematician, but as a literature detective.
To be scrupulously fair to the people who built it: OpenAI's technical paper on GPT-5's scientific contributions, written by serious mathematicians, described the work accurately. It claims four new results in mathematics, "carefully verified by the human authors," and calls the contributions modest in scope but profound in implication. Modest in scope. That is the company's own language, in the company's own paper. The paper was careful. The post was not. The distance between those two documents is the subject of this essay.
Because "solved" is doing a colossal amount of work in this field, and it means at least four different things. Tell them apart and most AI-science headlines decode themselves in about ten seconds:
Rediscovery. The system finds an existing published solution nobody had indexed. Genuinely useful — literature search at superhuman speed is a real service — but no new mathematics exists afterwards that did not exist before.
Assisted proof. A human mathematician works with the model, the model contributes real steps, and a human verifies the result. This is what OpenAI's paper documented.
Autonomous new result, independently verified. The machine produces something new; named experts check it and sign off. Rare. Real. We will get to it.
Reported but unverified. Somebody posts a claim and verification is described as "underway."
Category four is not hypothetical, and not a 2025 story. In April 2026, headlines announced that a newer OpenAI model had cracked a longstanding open Erdős problem in under two hours; the report was sourced to a social post, and noted that formal verification was underway. The headline ran first, the verification second. That is the order every time, and it is not an accident — it is what the incentives produce.
The business press did its part too, running the story as an AI that had stumped the world's best minds for decades. In fairness, the caveats were in that piece — that AI systems score far lower on open-ended research mathematics than on competition problems, that research-level autonomy remains elusive. They were around paragraph nine. Nobody shares paragraph nine.
So it was all just search, then? No — and that argument now has an expiry date
Here is where I have to argue against the version of me that felt clever in October 2025.
"It only searched the literature" was correct about that specific incident. As a general claim about AI and mathematics, it is now false, and anyone still using it is goalpost-shifting.
In May 2026, an OpenAI system disproved the Erdős unit distance conjecture — a problem that had stood for around eighty years. TechCrunch, to its credit, headlined the story with the memory of the last one: for real this time. And this time the company did the thing that was missing before. Alongside the announcement came published remarks by nine mathematicians — including the maintainer of the very database at the center of the earlier fiasco, the man who had called the previous claim a misrepresentation — presenting a short, human-verified digest of the counterexample. One of them, a Fields Medallist (who also co-authored OpenAI's own science paper, so not quite a disinterested outsider), called it a milestone; another called it the first autonomously produced AI result he found genuinely exciting. Why it worked is instructive: it played to the machine's strengths, grinding through hundreds of proof strategies a human would have abandoned after the first few.
That is category three. It counts. I am not going to pretend otherwise because it inconveniences my thesis.
And it is not isolated. Google DeepMind's AlphaEvolve was run across 67 problems in analysis, combinatorics, geometry and number theory, co-authored with mathematicians including Terence Tao; it rediscovered the best known solutions in most cases and improved on them in several. Better still, a subsequent Gemini case study addressed thirteen problems marked "Open" and did the arithmetic in public: five through apparently novel autonomous solutions, eight by identifying previous solutions already in the literature. Five and eight. Stated in the abstract, by the people with every incentive to blur it.
That is the honest version of the same exercise that blew up the previous autumn. Same task, same database, same category confusion available — and they simply reported the split. It cost them nothing. It made the paper better.
The most useful framing I found comes from Tao himself, writing on his own blog about the AlphaEvolve results: these systems tend to beat the state of the art where the literature on a problem is scant, and typically match or slightly underperform the known bounds where there is extensive prior work. Read that twice, because it explains both stories at once. Where humans have already been, the machine catches up to us. Where nobody bothered to look, it can get ahead. That is a real and valuable capability. It is also nothing like "the AI is doing science now."
The comparison that matters isn't another country. It's the claim that survived
Normally at this point I would take you to another jurisdiction — how Brussels handles it, what Tokyo tried. Not this time. The clarifying comparison here is between two claims made about AI and science: one that fell apart within days, and one now embedded in the foundations of biology.
In 2024, the Nobel Prize in Chemistry was awarded in two parts: half to David Baker for computational protein design, and half jointly to the two researchers behind AlphaFold2, the AI model that predicts a protein's three-dimensional structure from its amino acid sequence. Why did that claim survive scrutiny when the Erdős claim collapsed within days?
Not because AlphaFold was more impressive — though it was. Because of what shipped alongside it. The prize announcement notes that AlphaFold2 has been used to predict the structures of virtually all 200 million proteins researchers have identified, and that more than two million people from 190 countries have used it. Every one of those users is, functionally, an auditor. Structural biologists compared its predictions against crystallography they had already done, and found where it was confident and right, where it was confident and wrong, and where it flagged its own uncertainty. The claim was made in a form that made it falsifiable at planetary scale, by strangers, for free.
That is the whole difference. Not "AI did something amazing" versus "AI did something trivial." A verifiable artifact released to people with the means and the motive to attack it, versus a number in a post.
And notice what the good 2026 result did: it shipped the digest, signed by nine named humans including its loudest previous critic. It behaved like AlphaFold. The lesson was learned — expensively, publicly, and to the industry's genuine credit. My complaint is not that this is impossible. It is that it is optional.
Now imagine the version where nobody notices
First, the version that already happened: a doctoral paper reported that AI-assisted materials-science teams produced 44% more new materials, 39% more patents and 17% more product innovations. Irresistible numbers, and they traveled — before the paper collapsed it had been cited by the European Central Bank and the Congressional Research Service and praised by leading economists. Then MIT disavowed it, stating it had no confidence in the provenance, reliability or validity of the data. No finding of fabrication has been published against anyone, and I am not making one. The point is narrower and worse: a claim about AI's scientific productivity reached central bankers and congressional researchers before anybody had checked the underlying data.
Now scale that up. An executive order issued on November 24, 2025 launched the Genesis Mission, a national program promising to double the productivity and impact of American science and engineering within a decade. The order itself, as one business research group noted, carried no additional federal funds. On July 22, 2026 the White House announced more than $5 billion in federal commitments expanding it. Commitments, note — not appropriations, not money spent. Even the funding announcement sits one category away from the thing it sounds like.
So picture 2029. Every major research agency reports an "AI-attributed discovery" count, because a metric that goes up is politically irresistible, and grant renewals quietly key off it. Nobody has settled what qualifies — so a lab that used a model to find three overlooked 1974 papers logs three discoveries, a lab that produced one genuinely new verified result logs one, and on the dashboard they look like four. Journals, drowning in submissions, start accepting model-assisted verification of model-generated proofs. A drug program reorders its priorities around a synthesis route no human has independently reproduced, because reproducing it would cost eight months and the dashboard needs a number this quarter.
None of that requires a single bad actor. It only requires that announcing stays cheap and verifying stays expensive — which is exactly the situation we are in today.
What the careful people are saying
This is not a fringe worry, and the people raising it are largely not AI skeptics.
In June 2025, four months before the post that opened this piece, a physical scientist at RAND wrote that expert consensus on what counts as a scientific discovery by an AI "may prove more elusive than expected" — that the words useful and even novel are unexpectedly fuzzy in practice, and that we would be arguing about this for years in a messy middle ground where some experts call the same output groundbreaking and others call it obvious. He was right by October.
Arvind Narayanan, who directs Princeton's Center for Information Technology Policy, and his co-author Sayash Kapoor — whose research is specifically on why AI-based science so often fails to reproduce — argue for treating AI as normal technology: enormously consequential, but bounded by how fast organizations and institutions can actually absorb it, rather than a separate species arriving from outside history.
Tao's own position, in a compilation of his public remarks that he has reviewed, is the most quietly deflating thing I read all week: the near-term payoff is not the most powerful model on the hardest problem, but medium-powered tools scaling up mundane essential work — literature review above all. Which is, if you are keeping score, precisely what GPT-5 did in October 2025. It performed the most valuable near-term function available to it, brilliantly, and got embarrassed for it because somebody sold it as something else.
Then there is what the field itself predicted, with a date stamp you must not skip. In the largest survey of AI researchers ever conducted — 2,778 people who had published at the top AI venues — the estimate was a 10% chance by 2027 that unaided machines could accomplish every task better and more cheaply than human workers, rising to 50% by 2047. That survey was fielded in the autumn of 2023 and only cleared peer review in the Journal of Artificial Intelligence Research in October 2025. It is not what researchers think today, and it predates this generation of reasoning models. It is a snapshot of the field before the recent leap — and even then, the number was one in ten.
Set that beside what is being promised. Anthropic's chief executive wrote in a 2024 essay that AI-enabled biology would compress fifty to a hundred years of progress into five to ten — a prediction, clearly labeled as one, that now underwrites a great deal of capital. AI-for-science is a flagship product, with a Harvard physicist quoted estimating that a current model executes scientific projects about as well as a second-year graduate student. And, as one science journalist put it plainly, discovery is invoked by AI companies as a justification for their existence.
That is the machinery. Not villainy — machinery. When your civilizational justification is discovery, every ambiguous result faces enormous pressure to be announced as one.
What does this mean for you?
You are not going to verify an Erdős proof this weekend. Neither am I. But you will be handed a dozen of these claims a year, and you can grade them in under a minute.
Ask which of the four you are looking at. Rediscovery, assisted proof, verified autonomous result, or unverified report? Most coverage tells you if you read to the middle. If it doesn't say, assume the weakest one.
Look for the words "verification is underway." That is a claim published ahead of its evidence. It might be true. It is not yet known to be true, and those are different states.
Ask who checked, by name. "Nine mathematicians including the person who debunked our last claim" is a different animal from "our researchers found." Named human verifiers are the best signal in this entire field.
Ask whether you could check it. The AlphaFold test. Did they ship an artifact — a database, a proof, a dataset — that outsiders can attack? Or did they ship a number?
Check the date on any statistic, including the reassuring ones. The survey above is a 2023 forecast wearing a 2025 publication date. Perishability cuts both ways.
Forgive yourself for forwarding things. Then send the correction anyway, to the same people, with the same energy. It feels ridiculous. Do it. Somebody has to go first.
The lesson, as I see it
I still think what happened in October 2025 was, underneath the noise, quite wonderful. A machine read a substantial fraction of the mathematical literature and found solutions a careful human curator had missed. That is a real gift to a real field, and it deserved a good headline of its own.
Instead it got a better headline that was not true, and burned a little of the public's trust — trust the genuinely stunning May 2026 result then had to spend extra effort earning back, with nine signatures and a digest and a headline that had to say for real this time.
That is the cost of overclaiming, and it is not paid by the people who overclaim. It is paid by the next honest result. Every inflated announcement raises the evidence bar for whoever comes after, and the ones who suffer most are doing careful work with modest funding and no communications department.
My vote? Judge the claim by what shipped with it. Not the confidence of the phrasing, not the prestige of the lab, not how thrilling it felt in the first four seconds — which, I promise you, is exactly how long it took to get me. Ask what a stranger could check, and how soon. Everything that has survived scrutiny in this field, from protein structures to an eighty-year-old conjecture finally falling, answered that question the same way: here is the thing, go and test it.
The machines are getting genuinely, seriously good. Our habits for talking about them have not caught up. Only one of those two problems is hard to fix.
Human-aligned AI needs an honest scoreboard, and building one is a slow, unglamorous, deeply necessary job — which is the sort of thing the HAIA Foundation exists to do. If you would rather have the caveats than the thrill, subscribe — that is the whole offer.




