1 · Concept overview
Established An Earth-system digital twin is a continuously updated, observation-constrained simulation of the planet that users can interrogate with what-if questions. The term arrived from manufacturing, but the thing itself is an extension of numerical weather prediction: a physics-based model of atmosphere, ocean, land and ice, kept close to reality by data assimilation, run at resolutions fine enough that thunderstorms, mountain winds and city-scale heat are simulated rather than statistically approximated. The flagship is the European Commission’s Destination Earth programme, implemented by ECMWF, ESA and EUMETSAT, which launched its core platform in June 2024 with two priority twins: one for weather-induced extremes and one for climate-change adaptation.
Established The value claim runs through a long chain, and every link is a separate discipline. Observations from satellites, radiosondes, aircraft, buoys and ground stations are merged into a best estimate of the current state (assimilation); a global model advances that state; regional downscaling adds local detail; impact models translate hazard into flooded houses, stressed grids or crop loss; and finally a person or institution makes a decision — to evacuate, to pre-position cash, to close a surge barrier, to re-zone a floodplain. Forecast skill is measured obsessively at the middle of this chain and almost never at its end. This brief traces the whole chain and asks where the evidence actually stops.
Frontier This slot owns a joint question its neighbours each touch and none owns. Climate Engineering records that regional model disagreement is the genuine scientific dispute over deliberate intervention, and that models are the only instrument anyone has for evaluating it before doing it. Smart Cities documents a decade of urban instrumentation whose measured link from data to improved outcomes is close to empty, and names programme evaluation as the adjacency the field failed to import. Public Policy Foresight shows that foresight products are rarely scored against outcomes and that institutions consume forecasts without validating them. The joint question is whether a high-resolution, continuously assimilated simulation of the Earth system creates measurable public value at the point of local decision — or whether digital twins are about to repeat, at planetary scale and public expense, the evaluation failure the smart-city record already documents at city scale.
Frontier The short answer this brief supports: the physics and the computing are the strong links; the 2023–2025 machine-learning results are real and replicated; the delivered European system exists and runs; the impact-model layer is the weakest scientific link because exposure and vulnerability data lag hazard data by decades; and the decision layer is essentially unmeasured — no twin programme anywhere has published an outcome-scored trial of whether its products change decisions for the better.
Established A note on sourcing. This brief was commissioned in September 2026 from the Institute’s research base. Reading-list entries without links are cited from the bibliographic record rather than re-fetched, and claims are dated no later than early 2026 unless carried by a linked source.
2 · Current scientific position
Established Numerical weather prediction is one of the best-verified computational enterprises in existence. Bauer, Thorpe and Brunet’s 2015 survey documented what practitioners call the quiet revolution: useful forecast lead time has grown by roughly one day per decade for forty years. The gains came from better observations, variational assimilation, higher resolution and ensemble methods. ECMWF’s deterministic model has run at 9 km grid spacing since 2016, and its 51-member ensemble moved to 9 km in 2023. Global centres receive hundreds of millions of observation values daily and assimilate tens of millions; the ERA5 reanalysis is the reference reconstruction of the recent atmosphere since 1940 and the training corpus for essentially every machine-learning forecast model.
Established Between 2022 and 2025, machine-learning emulators matched and then beat physics-based models on standard forecast scores. NVIDIA’s FourCastNet (2022) first showed a network trained on ERA5 could produce credible global forecasts at a fraction of the cost; Huawei’s Pangu-Weather (Bi and colleagues, Nature, July 2023) reported deterministic upper-air scores exceeding ECMWF’s operational high-resolution forecast on reanalysis-based verification. Google DeepMind’s GraphCast (Lam and colleagues, Science, December 2023) outperformed the ECMWF deterministic system on around 90 percent of 1,380 verification targets. GenCast (Price and colleagues, Nature, 2024/2025) extended the result to probabilistic forecasting, beating the ECMWF ensemble on well over 90 percent of its targets out to 15 days. NeuralGCM (Kochkov and colleagues, Nature, July 2024) showed a differentiable physics-ML hybrid could hold its own from days to climate-length integrations, and Microsoft’s Aurora (2025) extended emulation to air quality and ocean waves. The results replicated across independent groups on a shared benchmark, WeatherBench 2.
Established The operational world moved almost immediately. ECMWF built its own emulator, the AIFS, and made a deterministic version operational in February 2025 — the first machine-learning model run as an official product by a major centre, with headline scores reported up to roughly 20 percent better on some medium-range measures. Two caveats are as established as the results. First, every one of these systems is trained on ERA5 and initialised from analyses produced by physics-based assimilation; the emulators sit on top of the physical machinery, they do not replace it. Second, standard scores reward smoothness: emulators tend to blur fine structure at longer leads, and verification of rare extremes — the events twins exist for — remains the contested edge rather than the settled centre.
Frontier Kilometre-scale global simulation works, for short periods, at great cost. The DYAMOND intercomparisons (from 2019) showed global storm-resolving models at 1–5 km are feasible for weeks to months. Regional convection-permitting models are further along: the UK Met Office’s 2.2 km climate projections demonstrated genuinely better representation of hourly rainfall extremes and the diurnal cycle than coarser models, which is the strongest measured argument that resolution buys decision-relevant realism for some variables. Destination Earth’s extremes twin runs a continuous global configuration at roughly 4.4 km with on-demand finer nests; its climate twin has produced multi-decadal projections at around 5 km on EuroHPC machines, including LUMI. What km-scale does not automatically fix is also documented: cloud microphysics, land-surface fluxes and ocean mixing remain parameterised, and a model that resolves the storm can still misplace it.
Frontier Destination Earth’s delivered state, as of early 2026, is a running system with an unproven user story. The programme (funded in the low hundreds of millions of euros under the Digital Europe Programme, on the programme’s own figures) concluded its first phase in June 2024 with the launch of the core service platform, data lake and the two priority twins, and entered a second phase running to mid-2026 aimed at operationalisation and user uptake. The system exists and runs; what its own public reporting does not yet contain is evidence of the kind the rest of this brief asks for: named decisions, made differently because the twin existed, with outcomes scored against a counterfactual. The honest description is a large, technically successful infrastructure whose demand side is still being constructed.
Established Where the chain has been audited link by link, the unglamorous links dominate the error budget. One correction to a global elevation dataset roughly tripled estimates of the population exposed to projected coastal flooding (Kulp and Strauss, 2019) — a larger change than a generation of sea-level modelling improvement, delivered by input data rather than simulation. Vousdoukas and colleagues (2020) showed that translating hazard into an adaptation decision — where raising European coastal defences pays — concentrates the answer in a minority of coastline where most damages cluster, an impact-model result in which the climate model is only one term. In infrastructure modelling, Hines, Cotilla-Sanchez and Blumsack showed that abstracted network models give actively misleading vulnerability rankings compared with physics-based ones, and Rinaldi’s interdependency taxonomy remains the standard statement of how coupled failures outrun single-sector models. The pattern generalises: resolution upgrades at the hazard end of the chain are routinely outweighed by data quality at the exposure end.
Established Forecast skill has an enormous measured literature; forecast value has a thin one. The cost-loss framework relating skill to value is decades old, and controlled evaluations are few: an anticipatory cash-transfer study around the 2020 Bangladesh floods (Pople and colleagues, cited by name) found households receiving forecast-triggered payments before the flood peak ate better and sold fewer assets. That is close to the entire genre of outcome-scored forecast-value trials. Institutions consume model output without auditing it: the UK National Audit Office found most major government spending is never robustly evaluated, and the energy-modelling record shows what unvalidated institutional forecasting looks like — the IEA’s World Energy Outlook underestimated solar deployment year after year, documented publicly, without institutional correction. Nothing guarantees twin programmes will escape this pattern; nothing in their governance yet requires them to.
3 · Frontier questions
Frontier Does kilometre-scale resolution buy skill on the variables decisions actually use? The regional convection-permitting record says yes for sub-daily rainfall extremes; the global km-scale record is too short and too expensive to have produced equivalent verification. The open question is not whether 4 km looks more realistic than 25 km — it does — but whether flood-relevant rainfall totals, wind gusts and heat at specific points verify better, by enough to change a warning threshold decision.
Frontier Can machine-learning systems forecast events outside their training distribution? Emulators trained on ERA5 have seen forty-odd years of weather; the events that matter most — record-shattering heat, category-breaking storms in a warming climate — are by construction rare or absent in training. Case studies through 2025 cut both ways: emulators have tracked some record events well and blurred others. A systematic out-of-sample verification standard does not yet exist, and until it does, physics models retain the argument that conservation laws extrapolate and learned statistics may not.
Frontier Can assimilation itself be learned? Prototypes reported by ECMWF and academic groups through 2025 attempt forecasts directly from observations, collapsing the assimilation-forecast pipeline into one learned system. As of early 2026 it is a promising research direction, not a demonstrated replacement.
Frontier Do downscaling relationships survive a changing climate? Statistical downscaling assumes the learned map from large scales to local scales is stationary; a warming climate is a standing violation of that assumption of unknown size. Dynamical downscaling avoids the assumption at compute cost. Which errors dominate for which decision variables is genuinely unsettled.
Speculative Are a twin’s what-if answers trustworthy where it matters? The interactive selling point of a twin is counterfactual interrogation — add this reservoir, close this barrier, change this land use. Verifying a counterfactual is philosophically harder than verifying a forecast, because the alternative world never happens. For interventions with historical analogues, hindcast tests exist; for novel interventions the twin’s answer inherits every structural error of the model, unaudited. Climate Engineering is the extreme case: the regional response to aerosol injection is exactly where models disagree.
Frontier Is the observing system that feeds the twins secure? Radiosonde networks have thinned in parts of the world, ocean profiling depends on continuously replenished Argo floats, and satellite continuity depends on agency budgets. Denial experiments show skill degrades measurably when major observation streams are withheld; the twin agenda assumes an observational foundation that is maintained by unglamorous, perpetually contested funding.
4 · Technological bottlenecks
Established Compute and energy bound the km-scale programme. A kilometre-scale coupled climate integration is an exascale workload; multi-decadal runs occupy top-tier machines for months and produce petabytes. Emulators relax inference cost dramatically but not the training-data cost, which still requires the physics pipeline.
Established Legacy code meets new hardware. Operational models are decades of Fortran optimised for CPUs; the machines are now GPU-heavy. Porting or rewriting is a multi-year programme at every centre and a real constraint on how fast km-scale becomes routine.
Established Exposure and vulnerability data are the weakest link in the chain the twins claim to serve. Hazard modelling has ERA5; impact modelling has no equivalent. Building footprints, elevations of thresholds, asset values, population by time of day, fragility curves — these are incomplete nearly everywhere and sparsest where risk is highest. The Kulp and Strauss result quantifies what one input-data correction does to a global impact estimate. No resolution upgrade compensates for not knowing what is in the floodplain.
Frontier Verification methodology lags the models. Point-based scores penalise a km-scale model for placing a real storm slightly off-grid (the double penalty), and standard metrics reward the blur that emulators produce. Scale-aware and neighbourhood verification exist but are not yet the currency of headline claims, which means headline claims systematically flatter the wrong properties.
Established Observation continuity is a bottleneck nobody owns. The twins are downstream of satellite programmes, radiosonde budgets and float deployments funded by dozens of states on independent cycles. A twin programme can buy compute; it cannot by itself buy the persistence of the global observing system.
5 · Research dependencies
Established Cloud and convection microphysics. At kilometre scale, deep convection is resolved but cloud droplet and ice processes are not; they remain parameterised closures whose errors set the ceiling on precipitation realism. Field-campaign-constrained microphysics is a research dependency no amount of compute removes.
Established Land-surface hydrology and urban physics. The decision variables — river stage, soil moisture, urban heat — live in component models whose observational constraint is far weaker than the atmosphere’s. Flood forecasting inherits errors from rainfall, terrain, soil and infrastructure representation in sequence.
Frontier Uncertainty quantification for learned systems. Physics ensembles have a theory of spread; emulator ensembles are calibrated empirically and can be overconfident precisely in unprecedented conditions. Trustworthy ML uncertainty is an open research field on which the twins’ probabilistic products depend.
Established Decision science. The evidence on how people interpret probabilistic information is directly load-bearing: Budescu and colleagues showed that lay readers systematically misread the IPCC’s calibrated likelihood language, interpreting “very likely” as far less probable than authors intended. Building a twin without building on this literature reproduces a known failure at the last link of the chain.
Frontier Evaluation methodology from programme evaluation. Quasi-experimental methods — difference-in-differences on staggered product rollouts, regression discontinuity on warning thresholds — are standard elsewhere in public policy and almost unused on forecast products. This is a methods-transfer dependency, not a discovery dependency.
6 · Required experiments
Frontier The decisive test is behavioural, not meteorological, and this brief ranks it first. The decisive test is a pre-registered paired trial in which matched decision units make real operational choices, one arm served by the twin chain and one by the incumbent products, and the arms are scored on decision outcomes rather than on forecast skill. Decision units could be districts triggering anticipatory cash transfers, reservoir operators, or municipal heat-plan activations; outcomes are losses, costs and response times, not correlation coefficients. The Bangladesh anticipatory-action study shows the design is feasible at household level. Nobody has scheduled or funded such a trial for any digital-twin programme, and until one runs, the public-value claim rests on the assumption that better skill upstream becomes better decisions downstream — the assumption the Smart Cities record specifically warns against.
Frontier Second: an out-of-distribution extremes gauntlet for emulators. A standing verification suite of record-breaking events held out of training — and, prospectively, each new record as it occurs — scored with scale-aware metrics against both emulators and physics models. Pieces exist in case-study form; the systematic version would settle the sharpest open question about the machine-learning results, and requires no new hardware.
Frontier Third: a resolution-value experiment on decision variables. Run the same events at 25 km, 9 km and 4 km through identical downscaling and impact layers, verify against dense observation networks, and publish where the added resolution changes the impact estimate enough to cross a real decision threshold. This isolates the link the programmes advertise from the links that dominate the error budget.
Frontier A natural experiment is already running and should be instrumented, not wasted. Destination Earth’s second phase is, in effect, a staggered rollout of twin services across European institutions. Staggered adoption is exactly the structure quasi-experimental evaluation needs; recording who adopted what and when, against outcome series, would turn programme administration into evidence at near-zero marginal cost. The window closes as adoption becomes universal or the programme is judged — either way — without it.
7 · Engineering requirements
Established The hard engineering is throughput under deadline. An operational twin must ingest global observations, assimilate, integrate and serve products inside fixed schedules — warnings that arrive after the flood are research, not service. This dictates the whole architecture: streaming pipelines, reserved compute, and graceful degradation when a data stream fails.
Frontier On-demand configurability is the genuinely new requirement. Classical NWP runs a fixed configuration on a fixed clock; a twin promises user-triggered nests, scenarios and sensitivity runs. Scheduling arbitrary interactive simulation against guaranteed operational deadlines on shared exascale machines is an unsolved resource-management problem being worked out in practice on the EuroHPC systems.
Established Data logistics outweigh model runtime. Kilometre-scale output cannot be stored exhaustively or moved casually; the engineering answer is compute-near-data, server-side reduction, and APIs that ship answers rather than fields. The Destination Earth data lake is an implementation of exactly this, and its usability — not its volume — is what user uptake will turn on.
Frontier A verification service belongs in the architecture. Skill monitoring is standard; what the decision-value question requires is an outcome-linkage layer — systematically joining issued products to downstream actions and losses. No operational centre currently runs one, and nothing about it is technically difficult; it is unbuilt because nobody’s mandate requires it.
8 · Adjacent technologies
Established Urban digital twins are the same wager at city scale. Smart Cities documents the measured record of urban instrumentation: capability real, situational awareness real, outcome evidence nearly absent, evaluation methodology conspicuously unused. Earth-system twins share vendors, rhetoric and the missing final link; the two fields will likely be judged together.
Frontier Deliberate intervention is the highest-stakes customer. Climate Engineering records that the regional consequences of solar radiation modification are the genuine scientific dispute, and that the outdoor experimental record is nearly empty. Any intervention decision would lean on exactly the counterfactual mode of a twin — which is why the counterfactual-validity question above matters beyond forecasting. Weather Modification carries the cautionary evaluation record at smaller scale: decades of seeding programmes whose effects took randomised designs and half a century to pin down.
Established Foresight institutions are the demand side. Public Policy Foresight shows governments consuming forward-looking products without scoring them, and documents the uncertainty-communication evidence this brief’s ethical section leans on. A twin is a foresight product with better physics; it inherits the same institutional pathologies unless deliberately exempted from them.
Established Impact chains and operational decisions already exist to learn from. Compound Climate Hazards covers the coupled-failure modelling a twin’s impact layer must get right, and the evidence that abstracted infrastructure models mislead. Coastal Defense Systems describes the Maeslant barrier, whose automated closure rule already lets a forecast drive an irreversible operational decision with published reliability analysis. Climate Migration Planning consumes the downscaled projections whose input-data sensitivities this brief documents.
9 · Institutional requirements
Established The producing institutions are strong; the evaluating institution does not exist. ECMWF, the national meteorological services, ESA and EUMETSAT constitute a mature, treaty-based production system with decades of operational discipline. There is no counterpart body whose mandate is to measure whether forecast products change outcomes. The UK National Audit Office finding — that most major spending is never robustly evaluated — describes the default any twin programme inherits unless its funders demand otherwise.
Established Open data policy built the machine-learning wave, and is not guaranteed. The 2023–2025 emulator results were possible because ERA5 and operational analyses were openly available; every headline model trained on European public data. As private forecast products proliferate, the incentive to enclose training data and verification access grows.
Frontier The public-private boundary is being redrawn. Technology firms now produce forecasts competitive with state centres at near-zero marginal cost, while depending on state-funded observations and assimilation. The division of labour — who runs the physics, who runs the emulators, who is liable for a blown warning — has not been negotiated: a national service issuing a warning from a third-party emulator owns a failure it cannot fully inspect.
Established The global South is the stated beneficiary and the thinnest link. The WMO’s Early Warnings for All initiative (launched 2022, targeting universal coverage by 2027) exists because roughly half of countries lacked adequate multi-hazard early-warning systems. Twins concentrate capability in institutions that already have it; the binding constraints elsewhere are observation networks, national hydromet capacity and last-mile communication, none of which a European twin supplies by existing.
Frontier Procurement could purchase the missing evidence cheaply. A funder that conditioned the next twin phase on a pre-registered outcome trial and an instrumented rollout would obtain, for a rounding error of the programme budget, the one result this field lacks. Nothing structural prevents this; nothing currently requires it.
10 · Ethical & societal considerations
Established Forecast skill is distributed unequally, and twins widen the gradient by default. Skill is highest, and improving fastest, where observations are dense and institutions are rich; mortality from weather extremes concentrates where both are thin. A programme that raises the ceiling without raising the floor increases the inequality of protection even while improving the global average.
Established Uncertainty communication is a measured field with uncomfortable results. Budescu and colleagues showed that the IPCC’s calibrated uncertainty language is systematically misread by lay audiences. The trust evidence is more encouraging: Kerr and colleagues found that communicating evidence transparently, uncertainty included, does not undermine public trust, and Schneider and colleagues found the effect of uncertainty communication on trust depends on how the message aligns with prior beliefs rather than on the uncertainty itself. The design implication is direct: a twin’s choices about how to show confidence are the part of the system the whole chain funnels through, and they can be tested.
Frontier Automation bias attaches to photorealistic simulation. A km-scale visualisation looks like a photograph of the future, and looking real is not being right where realism outruns skill. Doctrine for when a forecaster should overrule the twin — and evidence on whether they still feel able to — is thin.
Frontier The impact layer encodes whose losses count. Exposure databases undercount informal settlements; asset-value weighting ranks a marina above a slum; a twin optimised against recorded damages inherits the record’s omissions. These are data-governance choices presented as technical defaults, and they determine who the system protects.
Speculative Anticipatory action shifts the moral bookkeeping. Acting on a forecast means sometimes acting on a false alarm, and compensating action taken for events that never came is politically harder than compensating losses. If twins push institutions toward anticipatory spending — which the humanitarian evidence tentatively supports — the legitimacy machinery for wrong-but-reasonable action has to be built alongside.
11 · Civilizational implications
Frontier Adaptation is the century-scale customer. Trillions in infrastructure will be sited and specified against climate projections. If km-scale twins genuinely narrow local uncertainty — unproven, as this brief documents — the compounding value across decades of capital allocation is enormous; if they mainly add vividness to unchanged uncertainty, they will have made overconfident planning easier. The difference between those futures is precisely the unrun validation this brief keeps returning to.
Speculative A single authoritative simulation is a monoculture risk. Forecast diversity across independent centres is an error-catching mechanism with a long record; consolidation onto one twin platform, however good, couples everyone to its structural errors at once.
Established Weather services are among the best-documented public goods in existence. Benefit-cost estimates for basic meteorological infrastructure routinely run far above one. The civilizational case for the observation-assimilation-model commons is made; the open question is only whether the marginal euro now buys more value in resolution, in exposure data, or in the last mile.
12 · Timelines
These horizons track the chain from computational capability to demonstrated decision value, on current funding and institutional behaviour.
- 10 yr: Frontier Hybrid physics-ML pipelines standard at every major centre; km-scale global simulation routine for extremes on demand; learned assimilation prototypes operational somewhere; at least one outcome-scored decision trial published if any funder requires it — the trial is feasibility-limited by will, not technology.
- 25 yr: Speculative Continuous km-scale coupled twins with validated impact layers in rich regions; exposure data approaching parity with hazard data where states invest; forecast-value evidence either institutionalised (an evaluation body exists) or the field has plateaued as infrastructure without demonstrated marginal value, as urban instrumentation did.
- 50 yr: Speculative Sub-km global simulation feasible; the binding constraints are, on this brief’s analysis, unchanged in kind: observation continuity, exposure data, and institutional consumption of uncertainty. Twin outputs plausibly embedded in adaptation law and insurance by default.
- 100 / 250+ yr: Handwave A verified planetary management simulation — trusted for intervention decisions — is asserted in programme rhetoric far more often than any verification pathway to it is specified; nothing in the current record licenses a date.
13 · Technology tree & dependencies
- Depends on Nothing on this map blocks the modelling itself — the compute, physics and machine-learning results documented here arrived without waiting on any sibling brief. The value claim, however, depends on institutional results other briefs document as unsolved: the evaluation discipline whose absence Smart Cities measures at city scale, and the outcome-scored consumption of foresight products that Public Policy Foresight shows governments do not currently practise.
- Requires (not on this map) Microphysics and convection closures validated against dedicated field campaigns, because resolution exposes rather than removes them; global exposure and vulnerability data brought to parity with hazard data, since input corrections currently move impact estimates more than model upgrades do; an institutional standard that forecast products face pre-registered outcome trials, without which the decision-value claim stays unmeasured; industrial-scale reserved exascale allocations so operational twins are not guests on research machines; continuity of the satellite and in-situ observing networks every link downstream silently assumes; and paying demand for impact-layer services beyond what free national forecasts already provide, because the business case for the last mile is otherwise unfunded.
- Enables Anticipatory humanitarian and municipal action with quantified triggers; adaptation planning whose local uncertainty is narrowed rather than merely rendered; an auditable evidence base for any future deliberate-intervention debate; and — if the validation programme runs — the first demonstration in any instrumentation field that simulation upstream measurably improves decisions downstream, a result every other telemetry-and-model agenda on this map would inherit.
- Adjacent Climate Engineering as the highest-stakes prospective consumer of counterfactual simulation; Smart Cities as the same wager at city scale with the evaluation record already in; Public Policy Foresight for the demand-side institutions; Compound Climate Hazards for the coupled impact chains; Coastal Defense Systems for the rare existing case of forecast-driven irreversible operational decisions.
14 · Common misconceptions & speculative claims
Handwave “Higher resolution means better decisions.” This is the load-bearing assertion of the whole programme rhetoric, and it is exactly the step that works by assertion. Resolution demonstrably improves realism for some variables (sub-daily rainfall extremes are the measured case); decisions depend on the full chain, where input data corrections have moved impact estimates more than resolution upgrades have, and where no outcome-scored trial of any twin exists. The claim may be true; it is currently unmeasured, which is different from established.
Established “AI has made physics models obsolete.” False in a specific, checkable way: every headline emulator is trained on ERA5 and initialised from physics-based assimilation. The emulators replace the forecast integration step — genuinely, cheaply, and often with better scores — while depending on the physical machinery for their training data and starting states. Kill the physics pipeline and the emulators lose both their initial conditions and their ability to learn a changing climate.
Frontier “A digital twin is just the old model rebranded.” Half true, and the half matters. The physics, assimilation and much of the code are continuous with NWP; genuinely new are continuous km-scale operation, user-triggered configurability and the fused impact layer. Whether that increment justifies the budget is the open value question — but the increment is real.
Established “Twins will soon predict weather weeks or months ahead.” Chaos sets a hard ceiling: deterministic midlatitude weather prediction degrades toward uselessness around two weeks, a limit the emulators approach faster and cheaper but do not move. Beyond it lies probabilistic subseasonal skill for specific windows of opportunity, real but modest, and no twin changes that arithmetic.
Speculative “A twin of Earth lets us test geoengineering safely before doing it.” The claim inverts the evidence: the regional response to intervention is where climate models disagree with each other, as Climate Engineering documents, so the twin’s answer on exactly the decision-relevant question inherits the disagreement. Simulation is the only pre-deployment instrument available, which is an argument for the long verification programme, not for present trust.
Frontier “Public twin spending duplicates what technology firms now give away.” The free emulators are downstream of public observations, public assimilation and public reanalysis; the gift depends on the commons it appears to replace. The accounting that treats private forecasts as substitutes for the public pipeline double-counts the same infrastructure.