1 · Concept overview
Between 2022 and 2025 machine-learning models trained on a reanalysis of the atmosphere caught and then passed the physics-based systems that had held the forecasting record for forty years, on the scores those systems were optimised for. Established That happened, it is not in dispute, and it is a narrower achievement than the headlines carry. The models learned to advance an atmospheric state forward in time. They did not learn to determine what that state is, and the determination — data assimilation, running on observations from satellites, radiosondes, aircraft and buoys — is where most of the cost, most of the institutional apparatus and most of the accumulated skill of numerical weather prediction actually lives.
Frontier This brief owns forecast skill itself: what is measured, at what lead time, for which variable, against what reference, and where the measurement misleads. Earth-System Digital Twins owns the chain from observation to decision and the finding that nobody has measured its last link. That division is deliberate and this brief does not restate it: the question here is not whether a better forecast changes a decision but whether these forecasts are better, in which respects, and on evidence that survives contact with how they were scored.
Frontier The answer this brief lands is that the aggregate-score victory is real and the scoring rule flatters it, that the models remain downstream of the physics stack for their initial conditions, and that the two open questions — extremes and physical consistency — are the same question seen from different ends. A model trained to minimise squared error converges toward the conditional mean, which is smoother than the atmosphere; smoothing wins the metric and loses the gust, the sting jet and the deepest pressure. Established The field knows this, has named it, and is building ensembles and generative formulations specifically to escape it.
Established A note on sourcing. This brief was commissioned in September 2026 from the Institute’s research base. Reading-list entries without links are cited from the bibliographic record rather than re-fetched, and claims are dated no later than early 2026 unless carried by a linked source.
2 · Current scientific position
Established The baseline being beaten is one of the best-verified computational enterprises in existence. Useful forecast lead time grew by roughly a day per decade for four decades, through better observations, variational assimilation, higher resolution and ensembles. That record matters here for one reason: it means the comparison has an agreed scoreboard, maintained by centres with no interest in being beaten, which is a condition almost no other machine-learning benchmark satisfies.
Established The skill results are specific and they are wins. A graph neural network trained on reanalysis reported better scores than the operational high-resolution physics forecast on about 90 per cent of roughly 1,380 verification targets across variables, levels and lead times. Established A three-dimensional transformer reported comparable or better deterministic scores against the same operational baseline. A diffusion-based ensemble model reported better probabilistic skill than the operational ensemble on about 97 per cent of roughly 1,320 targets out to fifteen days, including tropical-cyclone track. Each of those numbers is a count of target-by-target comparisons and not an effect size; a model can win 97 per cent of targets by small margins and lose the ones that matter operationally.
Established The cost difference is the least contested fact in the field and the largest in magnitude. A ten-to-fifteen-day global forecast that occupies a supercomputer for the better part of an hour is produced by a trained network in seconds to minutes on a single accelerator, after a training run measured in weeks. That is a four-to-five order of magnitude change in the marginal cost of a forecast, and it is what makes large ensembles, rapid re-forecasting and on-demand scenario generation arithmetically available for the first time.
Established Operational adoption has happened, and it happened at the centre with the most to lose. The European centre put a machine-learning forecasting system into operations in early 2025 alongside its physics-based suite, not in place of it, and has since extended the approach to ensemble products. The published improvement figures are the centre’s own and are reported as headline-score gains over its physics model for several variables; treat them as programme figures pending independent verification against observations rather than against analyses.
Established Every operational machine-learning forecast in 2026 is initialised from an analysis produced by a physics-based assimilation system. The network replaces the forward integration; the observing system, the quality control, the bias correction and the variational or ensemble assimilation that turn tens of millions of daily observations into an initial state all remain. Frontier This is the single most important structural fact in the subject and it is routinely omitted: the machine-learning models have not made numerical weather prediction cheap, they have made one stage of it cheap, and that stage was never the expensive one institutionally.
Established The training corpus is a reanalysis, and a reanalysis is a model output. The reference reconstruction of the atmosphere since 1940 is itself produced by running a physics model with assimilation over the historical observing record, so the machine-learning systems are trained to reproduce a physics model’s best estimate of reality, inheriting its biases and the inhomogeneities introduced as the satellite record changed underneath it. Frontier Scores computed against that same reanalysis therefore reward agreement with the training distribution, which is a circularity the field acknowledges and has not fully removed from its headline numbers.
Established The smoothing problem is measured, named, and structural rather than incidental. A deterministic model trained on mean squared error is trained to output the conditional mean of the predictive distribution, which has less variance than any single realisation of the atmosphere. Established The consequence is visible in power spectra: machine-learning forecasts lose energy at small scales with increasing lead time faster than physics forecasts do, and published analyses of this limitation appeared alongside the first wave of skill claims rather than after them. Generative and ensemble formulations are the field’s answer, and they demonstrably restore spectral energy; whether they restore it in the right places is the open part.
Frontier On extremes the record is mixed in a specific and interpretable way: position is good, intensity is weak. Case evaluations of severe European windstorms found machine-learning models capturing synoptic structure and track competitively while underestimating the small-scale wind maxima that determine damage. Tropical-cyclone track errors are competitive or better; minimum central pressure and peak intensity are commonly under-forecast. Both patterns are what the smoothing argument predicts, which raises confidence that the cause is understood even where the fix is not.
Frontier Hybrid formulations are the most interesting result of the period and the least discussed. A differentiable dynamical core with learned sub-grid physics reported skill competitive with operational forecasts at one to fifteen days while remaining stable in multi-decade integrations, which no pure emulator has demonstrated. That combination — short-range competitiveness plus long-run stability — is the property that separates a weather model from a climate model, and it is the strongest evidence that the physics is doing work no amount of training data substitutes for.
Established Nowcasting is the one sub-problem where a generative model was judged better by working forecasters rather than by a metric. A deep generative model of radar was rated more useful than competing methods by a large panel of expert meteorologists in the great majority of cases, on precipitation nowcasts out to ninety minutes. Frontier It is the cleanest evidence in the field that machine-learning forecasts can be operationally preferable rather than merely score better, and it is on the shortest lead time and the smallest domain, which is the regime where physics-based models were always weakest.
3 · Frontier questions
Frontier The organising open question is whether a machine-learning system can produce its own initial conditions. End-to-end systems that go from raw observations to forecast without a physics-based analysis have been demonstrated and report lower skill than the operational chain. If that gap closes, the entire cost argument changes, because the expensive half of the enterprise becomes learnable; if it does not, machine-learning forecasting is permanently a cheap accelerator bolted to an expensive observing-and-assimilation system that somebody still has to fund.
Frontier The second is whether out-of-distribution performance is a real limit or an artefact of how the question is asked. A model trained on 1940–present has seen a warming atmosphere but not the one that is coming, and the strongest concrete instance of non-stationarity is measured: the Arctic has warmed nearly four times faster than the globe since 1979, so the regional distributions these models learned are shifting fastest exactly where the observing record is thinnest. Speculative Against that, unprecedented events in 2023–2025 were forecast competently by models that had never seen their like, which is evidence that the learned dynamics generalise further than the pessimistic reading allows.
Frontier The third is what the scorecard should be. Root-mean-square error against an analysis is the metric that built the field and the metric that rewards blurring; continuous ranked probability score, spread-skill consistency and threshold-based extreme-value measures answer different questions and rank models differently. A community benchmark now exists and is genuinely useful, but a benchmark maintained partly by the institutions being benchmarked is a governance arrangement as much as a scientific one.
Frontier The fourth is sub-seasonal and seasonal, where the claims are loudest and the evidence thinnest. Machine-learning systems have reported gains at weeks three to six and on seasonal anomalies, but skill at those ranges is low for everyone, sample sizes of independent events are small, and the difference between genuine signal and a well-tuned climatology is hard to establish with a few decades of cases. A seasonal claim validated on fewer than thirty independent seasons is a claim with fewer degrees of freedom than it appears to have.
Frontier The fifth is the handoff, and it is a human-factors question with measured literature attached. Warnings are issued by people who now receive a product they cannot interrogate physically. Work on communicating uncertainty finds that transparent presentation of evidence does not undermine public trust, and that the effect of uncertainty communication depends on whether the message is consistent with prior belief. Neither result was obtained on machine-generated forecasts, and whether a forecaster’s ability to override degrades when the underlying model is not physically inspectable is unmeasured.
4 · Technological bottlenecks
Established The binding constraint is the observing system, and it is the one nobody can train their way around. Both the physics chain and every machine-learning system depend on satellite radiances, radiosondes, aircraft reports and surface stations; a gap in that network degrades the analysis, and a degraded analysis degrades every forecast initialised from it. Observing-system continuity is a procurement and treaty matter rather than a research one, which means the field’s most important dependency sits outside its control.
Frontier Second, verification against analyses rather than observations is a methodological bottleneck that flatters everybody. Scoring a forecast against the analysis produced by the same system family embeds correlated error; scoring against independent station and satellite observations is harder, noisier and rarer. The gap between analysis-verified and observation-verified skill is the number that would settle several arguments at once, and it is not routinely published.
Frontier Third, the reanalysis that trains everything is a public good with no dedicated funding model. It is produced by one centre, costs a great deal to recompute, and improves discontinuously; every machine-learning model in the field inherits its version. A monoculture in training data means a common-mode failure: a systematic bias in the reanalysis becomes a systematic bias in every downstream model simultaneously, and the usual defence — model diversity — does not apply when the diversity is downstream of the shared corpus.
Frontier Fourth, accelerator supply is now a forecasting dependency. Training runs compete for the same hardware as commercial model development, and operational centres are not the highest bidder. Inference is cheap, which is the point; training and re-training on each reanalysis revision is not, and a national meteorological service without a training budget becomes a downloader of other people’s weights.
5 · Research dependencies
Frontier The result this brief most needs is a machine-learning assimilation system that is genuinely independent of a physics-based analysis. Partial results exist — learned assimilation, observation-space networks and end-to-end demonstrations — and all currently sit below the operational chain in skill. Until one matches it, every claim that machine learning has replaced numerical weather prediction is a claim about one stage described as though it were the whole.
Frontier Second, an observation-space verification standard with an agreed extreme-value component. The field has the instruments: threshold-exceedance scores, extremal dependence indices and spectral diagnostics are all standard in verification research. What does not exist is a convention that headline comparisons must report them, which is why the same models can be described as beating physics and as failing on extremes without either party misquoting anything.
Frontier Third, held-out event sets that are genuinely held out. A model trained on 1940–2021 and evaluated on 2022–2023 has a clean temporal split; a model whose architecture was selected by looking at those years does not, and architecture selection is rarely reported with the same rigour as data splits. The discipline that would fix this is pre-registration of evaluation protocols, which exists nowhere in the field.
Frontier Fourth, impact and exposure data adequate to the forecasts. Where impacts have been quantified, corrections to exposure data have moved the answer more than model improvements did — revised elevation data tripled the estimated global population exposed to coastal flooding, and defence-investment analyses turn on exposure assumptions rather than on hazard skill. A forecast improvement of a few hours is worth less than an exposure dataset that is wrong by a factor of three is costly.
6 · Required experiments
Frontier The decisive demonstration is an end-to-end machine-learning forecast system, initialised from raw observations with no physics-based analysis anywhere in the chain, evaluated against independent observations on extreme-value metrics. Frontier It settles the two questions this brief treats as central at once: whether the cost argument extends to the expensive half of the enterprise, and whether the extremes deficit is a property of the training objective or of the initial conditions. The components exist, the first end-to-end systems have been published at lower skill than operations, and the comparison needs no new instrument — only a protocol nobody has agreed.
Frontier The second experiment is a verification audit that could be run this year on archives that already exist: recompute the headline comparisons against station and satellite observations rather than against reanalysis, and publish the difference. The expected result is that machine-learning advantages narrow, because analysis verification rewards agreement with the training target, and that they narrow most for the small-scale variables where the smoothing argument bites. If they do not narrow, the circularity objection is answered empirically and should be dropped.
Frontier The third is a standing out-of-distribution gauntlet: a curated set of record-breaking events, strictly outside every model’s training window, scored on intensity rather than on position. Prospective scoring is the only version that is not contaminated, since any retrospective set can leak through architecture selection. The Institute’s digital-twin brief proposes the same instrument for the twin chain; this brief specifies the skill-side variant, which differs in scoring the forecast rather than the decision.
Frontier The fourth is a forecaster-in-the-loop trial. Give matched groups of operational forecasters machine-learning and physics-based guidance for the same events, with the provenance masked, and score the warnings they issue rather than the fields they were given. The nowcasting precedent shows expert judgement can be elicited rigorously at scale. The informative failure mode is over-trust: a forecaster who cannot inspect the physics may under-override a confidently wrong field, and nothing published measures that.
Frontier The fifth is a deliberate degradation study, and it is the cheapest test of the observing-system dependency. Withhold classes of observation — a satellite sounder, the radiosonde network over an ocean basin — from the assimilation feeding both chains and measure how fast each degrades. If the machine-learning chain degrades faster, its apparent skill is partly borrowed from an analysis quality it does not itself produce, which is the quantitative form of this brief’s central structural claim.
7 · Engineering requirements
Established The engineering requirement that changed is inference, and it changed by orders of magnitude. A global forecast that required a reserved partition of a national supercomputer now runs on one accelerator in minutes, which makes ensemble sizes of hundreds rather than tens affordable and makes re-forecasting on demand a routine operation rather than a campaign. Frontier The operational consequence is that the limiting resource moves from compute to data ingest and product dissemination, which are the parts of a forecasting centre that were never the bottleneck and are not sized for it.
Frontier Training is the new capital item and it is lumpy. A full training run occupies substantial accelerator capacity for weeks and must be repeated when the reanalysis is revised or the architecture changes, which converts a continuous operational cost into a periodic procurement. That shape favours large centres and vendors over national services, and the observable consequence would be national services running other organisations’ weights rather than their own models.
Frontier Ensemble and generative formulations are the engineering answer to smoothing and they carry their own requirement: calibration. An ensemble is useful only if its spread matches its error, and spread-skill consistency is a measurable property that generative models do not acquire automatically. A generative forecast that produces sharp, physically plausible, miscalibrated fields is worse for a warning decision than a blurred but honest one, and the sharpness is the part that impresses on inspection.
Frontier Provenance is an engineering requirement here for an unusual reason: the training corpus is versioned and the forecasts are archived for decades. Synthetic Data and Training Provenance owns the general problem; the weather-specific instance is that a forecast archive whose models were trained on different reanalysis versions is not a homogeneous record. Climatologies computed across such an archive would contain artefacts of model training history rather than of the atmosphere.
8 · Adjacent technologies
Established The direct neighbour is the twin chain, and the division of labour is explicit. Earth-System Digital Twins owns everything downstream of the forecast field: downscaling, impact modelling, dissemination and the unmeasured last link to a decision. Its central finding constrains this brief’s: skill improvements are an input to a chain whose output has never been scored, so a demonstrated skill gain is a necessary and demonstrably insufficient condition for public value.
Frontier The methodological neighbour is automated science. Artificial Scientists owns the question of what it means for a model to produce a correct result without a mechanism a person can inspect, and weather is the strongest case study available for it: a domain with a long-run scoreboard, an agreed ground truth and a trained profession whose job is to disagree with the model. If inspectable mechanism turns out not to matter where it can be tested most cleanly, that is evidence with reach well beyond meteorology.
Frontier On the application side the neighbours are the ones that consume the forecast. Infrastructure Resilience owns how networks respond to a warned hazard, and carries the relevant caution: models that look right structurally can misidentify which components actually fail. Compound Climate Hazards owns simultaneous extremes, which is where ensemble sharpness and calibration matter most, and Climate Health Adaptation owns heat-warning thresholds. Cheap ensembles make compound-event probabilities computable for the first time, which is the most valuable near-term use of the cost collapse and the least developed.
9 · Institutional requirements
Established The institutional novelty is that the best forecast models are now produced partly outside the meteorological services. The leading emulators came from commercial research laboratories, trained on a publicly funded reanalysis, and evaluated against publicly funded operational baselines. That is a working arrangement so far because the public inputs are open; it is fragile in exactly one place, which is that no instrument obliges the downstream models to remain open in return.
Frontier The second institutional question is who is liable for a machine-generated warning. National services carry statutory duties to warn, and those duties assume a forecaster who can explain the basis of a decision. A system whose output cannot be traced to a physical argument does not obviously satisfy an explanation duty, and no meteorological service has published a liability position on machine-learning guidance.
Frontier Third, evaluation discipline is the recurring institutional failure across this whole family of programmes. Audit bodies reviewing government spending have found evaluation routinely absent or too weak to support the claims made from it, and the forecasting case is not exempt: adoption decisions are being taken on score comparisons published by the adopting institutions. The remedy is unglamorous and cheap — independent observation-space verification, published on a schedule — and no body has been given the mandate.
10 · Ethical & societal considerations
Frontier The sharpest ethical exposure is an intensity bias in a life-safety product. If machine-learning guidance systematically under-forecasts peak wind and minimum pressure, warnings issued from it are systematically late or low, and the harm falls on the people least able to act on a marginal warning. The bias is documented in case evaluations and is exactly what the training objective predicts. Whether it survives in ensemble and generative formulations is the question a life-safety regulator should be asking and none has publicly asked.
Frontier The second is uncertainty communication, where the evidence is better than the practice. Transparent communication of evidence does not undermine public trust, and the effect of communicating uncertainty depends on whether the message aligns with what the audience already believes. Neither finding was obtained on machine-generated forecasts, and the temptation with a cheap, sharp, confident-looking product is to communicate the realisation rather than the distribution.
Frontier Third, the observing network is a global commons that the beneficiaries do not equally fund. Radiosonde coverage is thinnest across Africa and parts of the tropics, which degrades analyses everywhere and degrades local forecasts most. A technology that makes forecasting cheap while leaving observation expensive redistributes capability toward whoever already has the data, unless the saving is deliberately redirected into the network.
Speculative Fourth, the deskilling risk is real and unmeasured. If forecasters stop reasoning physically because the guidance no longer rewards it, the profession loses the capacity to recognise the case the model has never seen — which is precisely the case where human judgement was supposed to be the safeguard. This is an empirical claim with an obvious study design and no study.
11 · Civilizational implications
Established Weather prediction is the clearest demonstration available that learned models can outperform mechanistic ones on a scoreboard that predates them. The domain has an agreed ground truth, daily verification, a forty-year record of incremental progress and institutions with no incentive to overstate a competitor’s result. Whatever is concluded here about the relationship between skill and mechanism generalises further than most machine-learning results, because almost nowhere else is the test this clean.
Speculative The most consequential second-order effect is that forecast cost stops being a constraint on how forecasting is used. Hundred-member ensembles, event-specific re-forecasts, counterfactual runs and per-user tailored products all become arithmetically available, which shifts the discipline from producing a forecast to producing a distribution over scenarios. The constraint that replaces it is calibration, which is harder to verify and easier to fake.
Frontier The durable risk is a monoculture built on a single training corpus. Forty years of forecast improvement came from diverse modelling centres making different errors; a field in which every model is trained on one reanalysis has correlated errors by construction. The failure this predicts is not a bad forecast but a confidently agreed one, with the ensemble spread understating the error because every member inherited the same blind spot.
12 · Timelines
These horizons track what becomes verifiable about forecast skill, not what becomes announceable about it.
- 10 yr: Frontier Machine-learning systems are standard operational components at every major centre, run alongside rather than instead of physics models, and the first end-to-end observation-initialised systems are either competitive or demonstrably stuck. Expect observation-space verification with extreme-value components to become a publication norm, and expect machine-learning advantages to shrink when it does.
- 25 yr: Speculative Either learned assimilation matches variational and ensemble assimilation, in which case the physics-based forward model becomes a reference implementation rather than an operational one, or it does not, in which case the field settles permanently into a hybrid whose expensive half is unchanged. The discriminating observation is whether any centre retires its physics-based analysis.
- 50 yr: Speculative The training corpus problem becomes acute as the atmosphere moves further from the reanalysis distribution, and the field either maintains a continuously updated reanalysis as critical public infrastructure or accumulates a slow, hard-to-detect degradation in exactly the extreme regimes that matter.
- 100 / 250+ yr: Handwave Claims that learned models will extend useful deterministic prediction substantially beyond the classical predictability limit are coherent only as claims about better use of the initial state; the limit itself is a property of the atmosphere, and no training procedure changes it.
13 · Technology tree & dependencies
- Depends on nothing on this map for the modelling itself — the skill results arrived without waiting on any sibling brief. The value of those results depends on Earth-System Digital Twins, which owns the observation-to-decision chain and records that its last link has never been measured, and on the consumption side on Infrastructure Resilience and Compound Climate Hazards, which own what a warned hazard does to a network and to overlapping extremes.
- Requires (not on this map) an assimilation system learned rather than inherited, without which machine-learning forecasting remains one cheap stage inside an expensive enterprise; a verification convention that scores against observations and reports extreme-value measures, without which the same models can be truthfully described as winning and as failing; continuity of the satellite and in-situ observing network, which is a treaty and procurement matter outside the field’s control; a durable funding route for the reanalysis every model trains on, currently a public good with no dedicated line; a published liability position for warnings issued from machine-generated guidance, which no meteorological service has; and reserved accelerator capacity for operational centres that are not the highest bidder for it.
- Enables hundred-member ensembles at routine cost, event-specific re-forecasting and counterfactual runs, compound-hazard probabilities that were previously unaffordable to compute, tailored per-user products, and forecasting capability for services that could never afford to integrate a global model.
- Adjacent to Artificial Scientists on results without inspectable mechanism, Synthetic Data and Training Provenance on versioned training corpora inside long-lived archives, Climate Health Adaptation on warning thresholds, the Institute’s overshoot and lock-in slot on why a weather emulator is not a climate projection, and Weather Modification on the intervention question this brief does not touch.
14 · Common misconceptions & speculative claims
Handwave “AI has replaced numerical weather prediction.” Every operational machine-learning forecast in 2026 is initialised from an analysis produced by a physics-based assimilation system running on the conventional observing network. What has been replaced is the forward integration. The claim is not false about that stage and is false about the enterprise, and the difference is where the money and the institutions are.
Frontier “Beating the physics model on 90 per cent of targets means 90 per cent better forecasts.” Those figures count target-by-target comparisons across variables, levels and lead times, not effect sizes and not operational outcomes. A model can win almost every target by small margins on smooth fields and still lose on the small number of high-impact cases that a warning service exists to get right.
Frontier “The models cannot handle extremes because they have never seen them.” This is the right worry attached to the wrong mechanism. The documented deficit is intensity rather than occurrence, and it follows from a training objective that rewards the conditional mean, not from the rarity of the event as such. That distinction matters because it predicts a fix — generative and ensemble formulations — whereas the naive version predicts none.
Frontier “These are climate models now.” A weather emulator trained on a reanalysis reproduces the statistics of the observed period and has no mechanism guaranteeing conservation or stability outside it; the one architecture that has demonstrated multi-decade stable integration retained a differentiable physical dynamical core. Emulating a climate model and projecting a climate are different claims, and only the first has been demonstrated.
Frontier “Forecasting is now free.” Inference is nearly free; the observing system, assimilation, reanalysis production, training runs and dissemination are not, and they are the larger share. Treating the cheap stage as the whole cost is how an argument for cutting observing-network budgets gets made, which would degrade the analyses that every machine-learning forecast depends on.
Speculative “The skill gains are illusory because the scoring is circular.” This brief takes the circularity seriously and does not endorse this conclusion. Operational adoption at a centre with an interest in its own physics model, and the nowcasting result in which expert forecasters preferred generative output on an independent judgement task, are both evidence that survives the objection. The honest statement is that the margin is probably smaller than published and the direction is real.