The institution you probably trust most in daily life is a colored bar on your phone — and you have almost certainly never checked its work.
You have rearranged a week around that bar, maybe driven somewhere because it promised the snow would hold off until evening. On any app, in any country, the ritual is identical: open it, read the number, believe the number, plan the day.
What you probably never asked is what makes the number. Pressed, you might say "a supercomputer," in the tone of someone who has heard the word before. Somewhere there is a very large machine; it is doing physics; the physics is why the number is worth anything.
Last December the United States quietly changed the middle part of that sentence, and the way it changed is a better story than the headline anybody wrote.
Three models, one date
When the machinery underneath forecast offices changes, a numbered notice goes out that reads like a parts catalog. Effective December 17, 2025, with the 1200 UTC cycle, the National Centers for Environmental Prediction implemented three new models: the Artificial Intelligence Global Forecast System (AIGFS), the Artificial Intelligence Global Ensemble Forecast System (AIGEFS), and the Hybrid Global Ensemble Forecast System (HGEFS). Keep that date; secondary write-ups have moved it to January 2026.
The numbers are genuinely startling. AIGFS is the deterministic model: one run, one answer. NOAA says it delivers improved forecasts more quickly and efficiently, "using up to 99.7% less computing resources," than its traditional counterpart (mind the "up to"). The concrete version is easier to feel: a single 16-day forecast uses only 0.3% of the computing resources of the operational GFS and finishes in about 40 minutes. Administrator Neil Jacobs called it "a new paradigm for NOAA," delivering products faster and "at a lower cost."
The same release is candid about the tradeoff. AIGFS shows "a significant reduction in tropical cyclone track errors at longer lead times" and, in the next breath, "a degradation in tropical cyclone intensity forecasts." Better at where a hurricane is going. Worse at how hard it hits.
AIGEFS is an ensemble — 31 slightly different forecasts run together, because any honest forecast of next week is a spread rather than a point — at 9 percent of the operational ensemble's computing.
Then the third one, which is the real story. HGEFS is a 62-member grand ensemble: 31 members from the AI system and 31 from GEFSv12, the physics-based ensemble that has done this job for years. Half machine learning, half fluid dynamics, in one product. NOAA's claim about it is the cleanest sentence in the release — the hybrid "consistently outperforms both the GEFS and the AIGEFS across most major verification metrics." Read the next words, too: that is "initial testing," and "most major" metrics, not all. The world-first claim comes with its own caveat — "to our knowledge, NOAA is the first organization in the world to implement such a hybrid physical-AI ensemble system" — and work continues on its hurricane intensity forecasts.
So the headline is not that AI beat physics. It is that the thing forecasters were handed keeps the equations — and beat the two purebreds that went into it.
Where the skill comes from
Here is where I have to hold NOAA to its own text.
The summary bullet on the AI ensemble says "early results show improved performance over the traditional GEFS, extending forecast skill by an additional 18 to 24 hours" — the figure Fox Weather reported, attributing it to NOAA. The detail section says something quieter: "forecast skill is comparable to the operational GEFS." Improved, or comparable? The release does not resolve it, and I will not resolve it by picking the flattering half. Two further cautions about that number, since it is the one headed for a slide deck: it attaches to the AI-only ensemble, not the hybrid, and it describes forecast skill — how far out the output still verifies as useful — not how much warning you get before a tornado. Only one of those is measured in evacuations.
Where does a model get the ability to forecast a planet? By being shown one: a fine-tuned version of Google DeepMind's GraphCast, trained on NOAA's own Global Data Assimilation System analyses — the product of the physics-based system. A National Weather Service spokesperson, Erica Grow Cei, told CBS News that the latest models do not intend to replace the traditional ones, and that the traditional system is one of the sources the AI pulls from. NOAA's Daryl Kleist put it plainly in the same interview: "for these AI models, much of the gain in skill is owed to the fact that they were trained on analysis data." He adds a caveat that rarely survives retelling — those percentages describe making a forecast, and they "ignore the cost of training the models." The 0.3% is an inference number, not a lifecycle number.
The strongest case against me
Let me take the other side seriously, because it is stronger than I would like.
In August 2026, Google DeepMind's cyclone model was published in Nature with a result that is hard to wave away: its track, intensity and wind-radius predictions offer an average lead-time advantage of 1 day or more over leading operational models — "an improvement in accuracy comparable to the progress seen in the last decade of operational development," as the paper puts it. A decade of progress, in one model. Then read the authors' account of what it is for: "providing advanced operational ensemble guidance to human forecasters." Guidance. To forecasters. (No one in that sentence has been replaced.)
There is also a well-documented place where AI models still lose badly. A study published in Science Advances in April 2026 tested them on records and found that physics-based HRES consistently outperforms all AI models for hot and cold temperature records as well as wind speed records — even though the same AI models generally beat that physics model at ordinary 2-meter temperature across most lead times. They "generally underpredict temperature during high records and overpredict during low records," and "struggle when forecasting unprecedented events outside the training domain, even at short lead times." The researchers behind it, from the Karlsruhe Institute of Technology and the University of Geneva, are blunt about why. Zhongwei Zhang: "the greater the exceedance of the record of their training data, the larger the underestimation." Sebastian Engelke: physics-based models rest on fundamental laws, which is why their forecasts are still reliable when the atmosphere moves into states that have not yet been observed.
Sit with that. A model trained on the past is weakest precisely when the present stops resembling the past — the one condition a warming atmosphere keeps producing. Engelke's fix, reported by Carbon Brief, is "best of both worlds": physics for the unprecedented, AI for the cost. Which is a research paper describing what NOAA had already shipped.
Europe ran the same experiment in the open, and got the same answer
Is this just an American accident? Europe says no.
The European Centre for Medium-Range Weather Forecasts got there first, in public. On February 25, 2025, ECMWF took its Artificial Intelligence Forecasting System into operations — and the announcement holds the whole argument in one clause: to run "side by side with its traditional physics-based Integrated Forecasting System." Not instead of. Beside. The gains claimed are large: outperforming state-of-the-art physics models on many measures, "including tropical cyclone tracks, with gains of up to 20%," at roughly a thousandfold reduction in energy per forecast. ECMWF's director of forecasts, Florian Pappenberger, calls the pair "complementary."
On July 1, 2025, the ensemble version went operational with 51 members, disclosing the dependency NOAA discloses: it "relies on physics-based data assimilation to generate the initial conditions." ECMWF is "therefore also exploring hybrid systems that leverage the strengths of both approaches." Its scorecard is refreshingly two-sided — improvements up to 25 percent alongside degradations for forecasts of conditions higher up in the atmosphere. The peer-reviewed description in Geoscientific Model Development, June 1, 2026, reports overall skill improving by "4 %–6 %" in the upper air and near-surface variables. Impressive, and a long way from obsolete.
Then Europe did something I wish more agencies would copy: it turned the question into a contest. The AI Weather Quest is an open global competition ECMWF runs with the World Meteorological Organization co-sponsoring it, with teams submitting real-time sub-seasonal forecasts week after week. It kicked off in March 2025, divided into 13-week competitive forecasting periods, four in year one.
So who won? The team that post-processed the physics model. ECMWF announced on March 13, 2026 that the December-to-February winner was MicroEnsemble, whose approach uses AI to post-process state-of-the-art dynamical forecasts, from a field of 42 teams in year one. Its review of the season says it twice: that model "consistently led," outperforming the physics system for all variables and lead times, while the best data-driven and hybrid models "show limited significant improvements over the dynamical system, except for precipitation at both lead times" — and ECMWF says those initial results suggest that further development is needed before fully data-driven systems reliably and significantly outperform dynamical prediction models.
Two agencies, two continents, two different procedures — a verification suite in Maryland, an open tournament out of Reading (the town in England, not the verb). One answer: the winner keeps the physics and puts the machine on top.
The machine replaced nothing. The budget proposes to.
So why isn't this a feel-good story about hybrid vigor? Because of what the budget proposes.
AIGFS and AIGEFS were built by NCEP with NOAA's research laboratories and the Earth Prediction Innovation Center, inside the Office of Oceanic and Atmospheric Research — the research arm. NOAA's own FY2027 congressional justification proposes to terminate Weather Laboratories and Cooperative Institutes, about $95 million and 283 positions, closing named labs including the Global Systems Laboratory in Boulder and the National Severe Storms Laboratory in Norman, and discontinuing funding for the tornado field program VORTEX and the Warn on Forecast system. The request totals $4,540,824,000 — a decrease of $1,093,791,000 from the FY2026 enacted level.
Be precise: this is a request, not a law, and Congress has pushed back before. The Earth Prediction Innovation Center is itself funded in it, inside the Weather Service, so this is not "defund the AI." It is closing laboratories and moving what survives closer to operations — the administration's own rationale, as recorded by the Congressional Research Service: to "align research closer to operations" and focus on public safety.
At an April 28, 2026 hearing, Jacobs told lawmakers there was "really not any areas necessarily that would be abandoned. Some would be accelerated." The objection that stuck came from his own side: Rep. Brian Babin, the Texas Republican who chairs the House Science Committee, warned that eliminating the grants "would stymie future improvements." He added: "let us not forget that NOAA's primary mission is to protect lives and property."
Underneath the labs sits what the models eat: observations. On March 20, 2025, the Weather Service announced it would be stopping or reducing weather balloon operations at 11 locations because of staffing shortages. In July 2026, House Democrats wrote to NOAA alleging a cost — that there were no balloons launched upstream of eastern Kansas ahead of April 13, the day of the Kansas City tornadoes, and that no tornado watch was issued until 30 minutes before the first touchdown. That is an allegation in a members' letter, not an agency finding; hold it that loosely. Hold this tightly: a model trained on analyses is only as good as the observations behind them, and thinning the network to pay for the model is eating the seed corn with a very efficient fork.
The forecast in 2036
It is 2036. The hybrid ensemble stopped being news long ago. Your phone runs a model tuned to your postal code, updating every few minutes, and it is usually magnificent. It costs almost nothing to run, so three private companies run their own and sell the good one by subscription — roughly what the Project 2025 chapter on NOAA proposed when it said the Weather Service "should fully commercialize its forecasting operations." (A fact-check rated the harsher "get rid of NOAA" framing Half True; the commercialize line is a direct quotation.)
The failure mode I find plausible is not a robot uprising; it is a degradation nobody can see from outside. The models are still trained on analyses, and the analyses are still built from balloons, buoys, aircraft and satellites. Each year the network thins and the training set leans harder on what the model already believes. Skill scores stay excellent — they are computed against ordinary weather, which is what the model is best at. Then comes a day the atmosphere does something it has not done before, and the AI, weakest exactly where the record breaks hardest, smooths it toward the familiar. The physics model would have caught it, but by then it is legacy software maintained by three people.
And suppose the warning does arrive earlier. One of the few studies that asked — a stated-preference survey of 320 visitors to the National Weather Center in Norman, Oklahoma — found an average preferred tornado lead time of 34.3 minutes, and that given an hour, taking shelter became a lesser priority than when given a 15-minute lead time. Not observed behavior, not a national sample, and its authors say response "may be complex and situationally dependent." But it points at something no model can fix. Time is not warning. Time nobody uses is just a longer wait.
What the people who study this say
The scientists have been remarkably unanimous, and remarkably ignored.
A 2026 paper in the Bulletin of the American Meteorological Society argues that machine learning statistical estimation "neither can nor should be expected to supplant physics-based simulation". Researchers at NSF NCAR, whose own AI system flags severe-weather hazards days ahead, say flatly that the new forecasts don't negate the need for traditional, high-resolution weather modeling; their scientist Ryan Sobash puts it as "AI does not replace traditional models, but it can help us get more useful information in addition to those models."
The National Hurricane Center, asked in its own Q&A whether AI is coming for forecasters' jobs, answered: "we're learning that the answer is a resounding 'no'." It treated last season as an experimentation year and volunteers that there are "other examples where the traditional models performed better." Its science operations officer gives the logic in a sentence I would tape to every procurement officer's monitor — because the two families use different methodologies, "their forecasts have different types of errors, which can be very valuable to forecasters". Disagreement between your instruments is information. A single instrument cannot disagree with itself.
The question underneath is not accuracy but trust — which is why there is an entire NSF institute devoted to trustworthy AI for weather and coastal oceanography, directed by the University of Oklahoma's Amy McGovern, who holds appointments in meteorology and in computer science. When a field builds an institute around whether a tool can be relied on, that is the field telling you where the hard part is.
On the money the spectrum splits predictably, and agrees more than it admits. From the left, the Union of Concerned Scientists argues that weakening the integrated system means "less time to prepare, less certainty in forecasts" — putting the cut at 32 percent, or $1.6 billion, which counts differently from the request's own headline. From the right, Project 2025 wants forecasting commercialized, not abolished. In between sits Babin's line about lives and property, which is not a left or right sentence.
What does this mean for you?
Stop reading the icon and start reading the spread. Ensembles exist because the atmosphere is uncertain; the "60 percent" is the honest part and the little sun-behind-a-cloud is decoration. If your app hides probabilities, switch.
Find out what feeds your local forecast. Whether your region still launches balloons twice a day is a public fact, and so is your nearest forecast office's staffing. Both predict your warning quality better than any model's name.
Decide now what you would do with an extra hour. More lead time does not automatically produce more sheltering. Write the plan down, tell the household, pick the room.
When you see "AI beats weather model," ask three questions. Which model, which metric, which lead time? Ordinary temperature and record-breaking heat give opposite answers. And for decisions that matter — a roof, a harvest, a flight — read two sources, because forecasters value the two families together precisely for their different errors.
If you have a view on the labs, the record is open. The proposal names GSL, NSSL, VORTEX and Warn on Forecast on the page, and appropriations are decided by people who answer their mail.
The lesson, as I see it
I went looking for a story about a machine beating the physics, because that is the story I assumed I would find. What is on the record is stranger and more useful: a public agency built a machine-learning model that is dazzlingly cheap, tested it honestly, found it had borrowed much of its skill gain from the system it was meant to supersede — and shipped the crossbreed instead of the winner. Across an ocean, a different agency ran an open tournament and watched the same answer come back. That is not a technology failing to deliver; it is a technology placed correctly by people who verified before they deployed, and it got almost no coverage, because "we kept the old thing and added the new thing" is not a headline.
My vote? Grade the AI transition by what it is allowed to replace. A model that makes forecasters faster is a gift. A budget that treats the model as a reason to stop measuring the sky is an accounting trick with a long tail. The equations kept their seat on that ensemble because they earned it — and because somebody checked. Keep the checking.
A forecast is a machine's best guess about a future nobody has seen yet; what you do in the hours it buys you is the only part you control. If that distinction interests you, the HAIA Foundation has a whole publication made of it.





