1 · Concept overview
Civilizational planning is the claim that a society can adopt a plan whose horizon exceeds the careers of everyone who adopts it, and then execute it. The instances are real and dated: an atmospheric treaty signed in 1987 whose completion is projected for the 2060s, national five-year plans running since the 1950s, statutory duties toward people not yet born, repositories licensed to isolate waste for ten thousand years, and a global goal set with a hard deadline and scored against it.
This brief was commissioned to assess whether that capability exists. It gets part of the way and then stops, and the stopping point is narrow enough to name in one sentence. There is one plan at a genuinely civilizational horizon whose compliance is measured by a body with no stake in it and whose counterfactual has been modelled. There are two more with measured outturns and independent or ex-post verification, at horizons of roughly thirteen and roughly twenty years — short enough that section 2 rules on each explicitly rather than counting them silently. Everything else with a published outturn was scored by the planner. And the institutions built specifically to represent the future have been externally evaluated exactly once, by one audit office, in one country, without a counterfactual.
What is missing is specific rather than general. It is an evaluation that (a) takes a statutory long-horizon planning institution, (b) is run by a body with no stake in the answer, and (c) builds a counterfactual against outcome data rather than scoring process. That study has not been run, its natural experiment is sitting unused, and until it exists “the duty changed nothing” and “the duty prevented worse outcomes” remain observationally equivalent. What follows is the record, the measurement that does exist, and the exact shape of the hole.
2 · Current scientific position
Established One plan in the record has a civilizational horizon, an independent measurement system and a published counterfactual, and it is the ozone treaty. Total chlorine entering the stratosphere fell by 420 ± 20 ppt, about 11.5%, between the 1993 peak of 3,660 ppt and 2020; total bromine fell 3.2 ± 1.2 ppt, about 14.5%, from its 1999 peak. Projected return to 1980 ozone values runs around 2035 at Northern mid-latitudes, around 2040 near-globally, around 2045 in the Arctic and around 2066 over Antarctica — roughly eighty years from signature to completion, with the completion date itself a projection rather than an observation. The counterfactual is modelled rather than observed: holding CFCs growing at 3% a year from 1974 estimates about 1 °C of avoided global average warming and 3 to 4 °C in the Arctic by 2050, roughly a quarter of global warming mitigated as a side effect. Frontier It is the near-ideal case — few producers, cheap substitutes, an atmospheric signal that can be read in parts per trillion — and generalising from it is the step this brief refuses.
Established The compliance failure inside that treaty is the strongest available evidence that a long-horizon plan can be policed, and it is routinely told backwards. Unreported CFC-11 production was detected by the monitoring network, attributed to a region, and then stopped: the scientific assessment records that global CFC-11 emissions declined after 2018, with eastern China explaining 60 ± 30% of the decrease. The cost of the episode is bounded and the assessment gives both halves of it — a delay of about one year in mid-latitude equivalent effective stratospheric chlorine returning to 1980 levels, and up to three years at the poles. Detection, attribution, cessation, and a delay measured in single-digit years against an eighty-year plan is a verification system working, not a plan failing. Quoting only the three-year polar figure inverts what the source says.
Frontier Two further plans have measured outturns and verification by someone other than the planner, and this brief rules on both explicitly rather than leaving them for a reader to find. Smallpox eradication is ruled out on horizon, not on evidence. It was scored by the Global Commission for the Certification of Smallpox Eradication in 1979 — an independent verification body, which is precisely the property this brief says exists only for the ozone treaty, and the claim has to be corrected accordingly. Its intensified programme ran from 1967 to the early 1980s: roughly thirteen years, inside one career, which is the honest reason to exclude it from a civilizational count rather than any weakness in the verification. Frontier The United States Acid Rain Program is ruled in as a generational case. Title IV of the 1990 Clean Air Act Amendments produced a compliance outturn that was measured continuously rather than self-reported, and a substantial ex-post economics literature exists on it — Schmalensee and Stavins in the Journal of Economic Perspectives in 2013 is the standard entry point. Its horizon is roughly twenty years. Established So the honest count is one plan at civilizational horizon with independent measurement and a counterfactual, and two more with measured outturns and outside verification at horizons of roughly thirteen and roughly twenty years — not “exactly one”, which is the count this brief was commissioned with and is correcting.
Established Where a whole-society plan met a published deadline, the scorer was almost always the planner. The Millennium Development Goals are the cleanest global plan with a hard date and a published scorecard, and the scorecard is mixed. Extreme poverty: target halve, outturn 47% in 1990 to 14% in 2015, exceeded. Child mortality: target a two-thirds reduction, outturn 90 to 43 deaths per 1,000 live births — more than half, about 52%, missed. Maternal mortality: target three-quarters, outturn 45%, missed widely. Universal primary enrolment: 91%, not achieved. Established The successor goals are behind: on the United Nations' own 2025 accounting, only 35% of targets are on track or making moderate progress, nearly half are moving too slowly, and 18% have regressed, with five years remaining. Both scorecards are published by the body that set the targets.
Established China's five-year plans are the most-cited evidence of planning capability, and their attainment figures are endogenous by the admission of the people who collect them. The record on its face is strong: 20 of 23 metrics in the 11th plan, all 28 in the 12th, 24 of 33 met a year early in the 13th. Written testimony to the US–China Economic and Security Review Commission by Oliver Melton of the State Department, delivered 22 April 2015, records that “local governments and ministries are almost always responsible for collecting the data used to evaluate their own performance, which leads to misrepresentation and obfuscation”, and that favoured candidates “can also receive plum assignments where targets are easier to meet, or they can negotiate friendlier evaluation criteria at the outset”. Frontier A think-tank analysis reaches the same place from the opposite direction: high attainment is the start of the conversation rather than the end, because targets may simply have been set achievable. No independent body scores these plans. There is no audit finding, no external verification and no counterfactual, so attainment is evidence about a target-setting system rather than about whether the plan caused the outcome.
Established Exactly one institution built to represent the future has been examined at scale by an audit office, and the examination found it was not working as intended. Wales's Well-being of Future Generations Act 2015 has been reviewed twice by the Auditor General. The 2020 report covered 71 examinations across all 44 public bodies named in the Act and found bodies could demonstrate they were applying the sustainable development principle, with improvement needed across all five ways of working, “limited examples of forecasting and scenario planning”, and involvement practice focused on informing rather than involving. The 2025 report drew on approximately 200 separate audit reports over five years, May 2020 to May 2025, covering the 44 original bodies plus four Corporate Joint Committees and eight bodies added in June 2024. Its headline: “the Act is not driving the system-wide change that was intended.” Prominence has risen and conversations have changed; bodies “often did not have the right performance measures in place” and were “typically more focused on activities and outputs than outcomes”.
Established The decisive detail is what the auditor had to recommend in year ten. Recommendation R1 of the 2025 report asks the Welsh Government to “clearly set out a scope and timetable for its own post-legislative evaluation” — a decade after the Act, none had been carried out. The Welsh Government's written statement of 17 July 2025 said it “recognise[s] the call for post-legislative evaluation of the Act” and would consider a first-stage evaluation by an ESRC policy fellow and a Senedd committee inquiry when they conclude. Frontier That is the live variable for this slot and this brief could not resolve it. The Sixth Senedd has since ended, so the committee inquiry has either reported or lapsed, and the ESRC first-stage work was pointed at autumn 2025. If either has landed with outcome attribution, that plus the ozone treaty is two cases and this slot becomes arguable. The position stated here is the verified one as of 17 July 2025. Established Audit Wales is also not the only external body to have looked: the Senedd's public accounts committee published a scrutiny of the Commissioner's 2022-23 accounts on 5 March 2024.
Established And the outcome data that any causal claim would rest on is not built to carry one. The Welsh Government's own quality report for the national indicators states that progress is assessed only as “improved, deteriorated, no change or not assessed” and that “we have not considered whether the milestones are on course to be met, simply the direction of change”. There is no attribution to the Act anywhere in the statistical framework. The picture is mixed on the Commissioner's own account: the 30% volunteering milestone was hit and adult and child healthy-behaviour measures improved, while biodiversity declined, low birth weights reached their highest recorded level, youth participation in education and employment fell, and greenhouse gas emissions rose 7% between 2020 and 2021. Established On resources, the April 2022 Senedd scrutiny report gives the Future Generations Commissioner £1.509m for 2022-23 against £3.327m for the Welsh Language Commissioner and £5.287m for the Public Services Ombudsman. It is the lowest-funded commissioner in the country by £71,000 — the Children's Commissioner sits at £1.580m and the Older People's at £1.589m, both within 6%. The claim survives; the rhetorical framing usually built on it does not.
3 · Frontier questions
Established The first question is what kind of evidence the foresight-evaluation literature actually produces, and the answer is not “none”. There is an established sub-field: Georghiou and Keenan on evaluating national foresight activities (2006); Georghiou, Keenan and Miles (2010), written directly on the UK programme; Havas, Schartinger and Weber on foresight's impact on innovation policy-making (2010); Dursun, Türe and Daim (2011); Seidl da Fonseca (2016). What that literature evaluates is process, network formation, participant self-reported use and behavioural additionality — not decision counterfactuals. The distinction matters because the two claims sound alike and only one is true: it is false that nobody has evaluated foresight, and true that nobody located here has scored a decision against what would have happened without the foresight product. Established A count of “four evaluations, four finding no impact” is falsifiable and wrong, and it is corrected here rather than repeated.
Established What the official record does say, repeatedly, is that the products are good and the uptake is not measured. The Public Administration Select Committee reported on strategic thinking in government in April 2012, and that report triggered the review that followed. The Day review — January 2013, by Jon Day, Chairman of the Joint Intelligence Committee, published by the Cabinet Office — found “weak integration between Departments and with the policy agenda, particularly at the senior level”, that horizon-scanning products are “often lengthy, and poorly presented, making them harder to digest and easier to ignore”, and the load-bearing sentence: “previous reviews have attempted to embed cross-cutting horizon scanning into government structures without enduring success.” Established The Commons Science and Technology Committee in 2014 recorded that the Day review had found “an overabundance of reports [delivering] little in the way of policy change”, praised Foresight's outputs as of “impeccable quality”, and then recorded the Cabinet Secretary conceding that this work had “not always translated into actual policy changes”.
Established The study commissioned by the UK's own foresight body to find the impact evidence reported that the evidence base is limited. Eight country case studies — Canada, Finland, Malaysia, Netherlands, New Zealand, Singapore, the UAE and the United States — concluded that “there is a limited evidence base on the impact of foresight work. The majority of case studies that are available focus on how specific projects or units have used foresight rather than how governments as a whole have done this.” Its sharpest finding is a natural experiment: “despite pandemics being identified as a key issue in many foresight and other planning exercises, there was a failure to integrate, act or sustain attention”. Established The statutory verdict on that experiment is the highest-standing evaluation available anywhere in this subject: the UK Covid-19 Inquiry's Module 1 report of 18 July 2024 concluded the UK was “ill prepared for dealing with a catastrophic emergency” and had “prepared for the wrong pandemic”, with the 2016 Exercise Cygnus finding that preparedness was “currently not sufficient” not adequately acted on. Ten recommendations, groupthink among advisers, and a “labyrinthine” institutional structure. Frontier That record still does not separate anticipation does not work from anticipation works and governments ignore it, and the tie is inherited by this brief from Public Policy Foresight, which owns it.
Frontier The second question is whether the mechanism works at all under controlled conditions, and here the pack this brief was built from was wrong to record an absence. Japan's Future Design programme runs controlled experiments on exactly this mechanism. Kamijo, Komiya, Mifune and Saijo, in Sustainability Science in 2017, report that 60% of participants chose the sustainable option when a member of the group argued as an imaginary future generation, against 28% without. Related work by Saijo and colleagues has continued since. Frontier This is deliberative-mechanism evidence, not institutional evidence. It shows that a procedural device changes choices in a laboratory or a municipal workshop. It does not show that a statutory duty changes a government's decisions over decades, and nothing in it converts the missing institutional evaluation into an existing one. Saying “nobody has run it” is wrong; saying “it has been run at the deliberative level and not at the institutional level” is the accurate statement.
Established The third question is the denominator, and the best comparative study says the data has not been collected. Michael Rose's analysis in Politics and Governance identifies 25 institutional proxy representatives of future generations across 17 countries, of which 16 still exist in 11 countries. His effectiveness finding is that proxies are usually equipped to voice the construed interests of future generations but “often cannot act as watchdogs with teeth” when ignored, and that of the extreme cases — both the weakest and the strongest — “all have been dismantled”. His finding about the evidence is the one that matters here: he codes formal design features because that is what exists, and states that “the proxies' formal design features can only tell parts of their stories. Future empirical research should gather reliable comparable data on their resources” and substantive impact. In the field's leading comparative study, the impact data has not been collected.
Speculative The fourth question is the live hypothesis this brief refuses to promote: that political durability, not engineering, is the binding constraint on multi-decade programmes. The comparison usually offered is a repository terminated in the United States against one approaching an operating licence in Finland, with a municipal council vote of 20–7 and a parliamentary ratification of 159–3 on the Finnish side. It is a hypothesis and this brief flags it as one. The audit finding says only that the American termination was made “for policy reasons, not technical or safety reasons”; it names no legislator and identifies no veto player. Handwave The claim that the technical problems were comparable is supported by nothing cited anywhere in this brief's sources — unsaturated oxidising tuff assessed against a court-imposed million-year standard is not obviously the same problem as saturated crystalline bedrock with copper canisters, and no source establishes that it is. Established A brief that refuses to generalise from one ozone treaty cannot call an n=2 comparison “the measured variable”, and the correct status of the durability hypothesis is: plausible, consequential, untested, and lacking the denominator that would test it.
4 · Technological bottlenecks
Established Any plan rests on a projection, and the projection side of this subject is the one part with real measurement. It does not go well. Tetlock's twenty-year study accumulated 82,361 forecasts from 284 experts; as summarised in the International Journal of Forecasting, the experts “barely if at all outperformed informed non-experts”, neither group “did well against simple rules and models”, and on trichotomised outcomes the forecasters frequently did worse than assuming each state equally likely. Established Project forecasting is worse and better documented: nine in ten megaprojects overrun, rail averaging 44.7% cost overrun with 51.4% demand shortfall, roads 20.4%, and one in six ICT projects becoming an outlier averaging 200% overrun.
Established The finding that bears hardest on the framing under test is not the size of the errors but their constancy. Flyvbjerg's series reports that overruns “have stayed high and constant for the 70-year period for which comparable data exist”. Seventy years of expensive, visible, extensively documented failure produced no measurable improvement in forecasting accuracy. Speculative An institution that cannot improve its five-year forecasts across seventy years of feedback has no demonstrated mechanism by which it would improve its five-hundred-year ones — which is an inference about learning rather than a measurement, and is flagged as such.
Frontier Technology forecasting has a documented directional bias, and the analysis establishing it comes from an interested party. An ex-post study by Lopez, Keiner and Breyer at LUT University examined every World Energy Outlook published from 1993 to 2022 and reports systematic underestimation of solar photovoltaics across the whole thirty-year run: the 2022 net-zero scenario projecting 15.5 TW of solar by 2050 against the authors' own assessment of a feasible 63.4 TW, and a most-ambitious annual-installation peak of 657 GW a year in 2040 against markets expected to reach 1 TW a year between 2025 and 2030. The authors advocate 100%-renewable scenarios and the counterfactual is their own model, so the defensible claim is the direction and persistence of the bias, not its magnitude.
Established The measurement bottleneck on the planning side is structural, not incidental: the outcome data is not built for attribution. Wales assesses its national indicators as direction of change only and explicitly declines to ask whether milestones are on course. China's plan attainment is scored by the bodies being scored. The Millennium and Sustainable Development Goal scorecards are published by the body that set the targets. In each case the number exists and the counterfactual does not, and a number without a counterfactual cannot answer the question this brief was asked. Frontier That is a solvable problem in at least one case — the natural experiment described in section 6 requires no new method and no new data collection.
5 · Research dependencies
Established Nothing on this map produces a result this brief waits on, and no typed depends-on edge is claimed. Civilizational planning does not wait on a discovery. It waits on four measurements that institutions have chosen not to perform, all of them recorded as typed constraints in section 13: an independent evaluator with no stake in the answer, a counterfactual design applied to a statutory long-horizon institution, external verification of national plan attainment, and a survival denominator for multi-decade commitments. Established Each is something a legislature, an audit office or a research council could choose to supply. None requires new method.
Frontier Two adjacent briefs own parts of this question and are not duplicated here. Public Policy Foresight owns the machinery for anticipating the long term and the evaluation gap around it; Long-Term Institutions owns statutory duties, commissioner mortality, discount-rate guidance and repository marker design as institutional questions. This brief owns the plans themselves — documents with dates, targets and outturns — and the question of whether any of them has been independently scored.
6 · Required experiments
Established The single highest-value experiment in this subject is already set up and nobody has run it. Wales legislated a statutory future-generations duty in 2015 and the rest of the United Kingdom did not. That is a natural experiment with a control group and roughly eleven years of indicator data collected on a comparable statistical basis. At the institutional level it remains unrun. Running it requires no new instrument, no new survey and no cooperation from the body being evaluated — only a difference-in-differences design against outcome series that already exist, executed by someone with no stake in the answer. Frontier Until it is run, this slot cannot be closed, and that is the whole reason for its status.
Established Second: decision-tracing on foresight outputs, which is the specific thing the existing evaluation literature does not do. Not “was a report produced” and not “did participants report using it”, but: take N foresight products, trace forward to the specific decisions they were meant to inform, and score whether the decision differed. The UK's own commissioned study says the available case studies do the former; the 2026 Policy and Society paper says its authors did not attempt the latter because it “would have required further primary data collection and process-tracing” they did not undertake. Both statements identify the same missing study.
Frontier Third: independent scoring of national plan attainment. An external audit of target attainment — of the kind national audit offices routinely perform on individual programmes — does not exist for any multi-decade national plan located here. The data to do it is largely published; what is missing is an auditor with a mandate and no stake. Speculative Fourth: survival statistics with a denominator. One terminated repository programme is one observation and one licensed programme is another. The missing quantity is how many multi-decade statutory commitments have been made, how many survived a change of government, and what predicted survival. Rose's finding that both the weakest and the strongest proxies were dismantled is the beginning of such an analysis, on a sample of 25 with no resource or duration data attached.
Frontier Fifth, and the one with a running head start: scale the Future Design manipulation from the workshop to the statute. The 60%-against-28% result is a clean randomised effect on a deliberative choice. Whether the same device changes an appropriations decision, a planning consent or a capital programme when embedded in a standing institution is untested, and it is testable — the treatment is cheap, the assignment can be randomised across comparable local authorities, and the outcome is a decision rather than a survey response.
7 · Engineering requirements
Established The only bodies compelled to document decisions on civilizational timescales are nuclear waste programmes, and what they have actually measured is political durability rather than geology. The American case is the clearest measured failure of long-horizon institutional commitment on record. The Nuclear Waste Policy Act passed in 1982; Congress narrowed the search to Yucca Mountain in 1987; the Department of Energy missed its statutory deadline to begin accepting commercial spent fuel on 31 January 1998; it recommended the site in 2002, filed a licence application in June 2008, announced termination in March 2009, and the programme was dismantled on 30 September 2010. Total spend: nearly $15bn in constant FY2010 dollars, with $41–67bn more required to finish. The audit finding on why is unambiguous and narrow: the decision was “made for policy reasons, not technical or safety reasons”, the Secretary's judgment being “not that Yucca Mountain is unsafe… but rather that it is not a workable option”.
Established The bill is still running, which is the part that makes it a planning finding rather than a policy anecdote. As of 2020 the federal government had paid about $8.6bn in damages to reactor owners for the 1998 default, with total liability projected at $39.2bn; the Nuclear Waste Fund holds about $43bn and has taken no new receipts since the fee was zeroed on 16 May 2014; and roughly 86,000 metric tons of commercial spent fuel sits at 75 reactor sites in 33 states, growing by about 2,000 tons a year. More than forty years after the Act there is no repository. The audit office's stated lesson is that such programmes require “consistent policy, funding, and leadership, especially since the process will likely take decades”.
Established The institutional-memory case is a facility licensed for ten thousand years that was defeated at year fifteen by a procurement substitution. The accident investigation traces the radiological release of 14 February 2014 to a single drum and finds the local root cause was a laboratory substituting “organic, wheat-based absorbent instead of the directed inorganic absorbent such as kitty litter/zeolite clay”, under a repackaging procedure whose own wording specified organic absorbent where inorganic was meant. The systemic root cause was oversight failure by two federal offices, in a safety culture where workers “did not feel comfortable identifying issues that may adversely affect management direction”. The board's conclusion: “the release from the container(s) was preventable.” Emplacement stopped in February 2014, authorisation to resume came in December 2016, emplacement restarted in January 2017 and regular shipments in April 2017 — a three-year outage produced by an ambiguous sentence in a procedure written at another site.
Established The marker programme is the most honest document in this literature and its own caveats are the finding. The 1993 Sandia report convened a markers panel working as two teams, of six and seven members — materials scientists, architects, anthropologists, linguists, astronomers, semioticians — to design a warning system interpretable across the ten-thousand-year regulatory period, alongside a separate futures panel. The teams disagreed on fundamentals: one designed a site with no discernible centre relying on written redundancy and human facial expression, the other drew visitors toward a central information structure and leaned on pictographs. One team conceded that “evolution of existing cultures and the creation of new ones over the next 10,000 years cannot be known” and had to assume that “scholarship capable of translating the messages on the markers will continue to exist somewhere in the world”. Established The report's statement that such measures “can never be assumed to eliminate the chance of inadvertent and intermittent human intrusion” is about passive institutional controls generally, not about markers specifically, and quoting it as a verdict on markers overstates it. Frontier No marker system has been built — markers are erected at closure — so none of it has been tested against anything.
Frontier The counter-case took forty-three years and is not yet complete. Finland's government approved deep disposal by Decision-in-Principle on 10 November 1983; preliminary site characterisation ran 1987–1992 and detailed characterisation 1993–1999; Olkiluoto was selected in 1999; the host municipality's council voted 20–7 in 2000 and Parliament ratified 159–3 in May 2001; excavation began in 2004. The operator subsequently applied for an operating licence — this brief prints no date for that application, because the only source located is an industry-forum item dated 12 January 2022 that says merely “recently”, and publishing a publication date as an event date is exactly the defect this site's sourcing rule exists to prevent. The regulator issued a positive safety assessment on 4 August 2026, evaluating long-term safety “considering a time period of at least 100,000 years and as long as 1 million years”, with the government decision pending and a licence sought through 2070. Frontier Both the timeline and the votes come from the operator and from the nuclear trade press, which are interested parties, and the dates are the part of their account that is independently checkable.
8 · Adjacent technologies
Within this map: Public Policy Foresight, which owns the anticipation machinery and is itself published as ongoing research for a related reason; Long-Term Institutions, which owns statutory duties, discount-rate guidance and marker design as institutional objects; Existential Risk Governance, which inherits this brief's measurement problem at the tail of the distribution; Civilization Resilience Planning, where planning meets preparedness; Megaproject Governance, which owns the overrun series used here as a bottleneck; Nuclear Waste Solutions, which owns the geology this brief deliberately does not adjudicate; Future Public Administration, whose finding that evaluation is discretionary explains the base rate below; and Global Cooperation Models, which owns the treaty machinery the ozone case runs on.
Outside it: atmospheric chemistry and the ozone assessment process, which supplies the only independent long-horizon measurement here; the forecasting and judgement literature; welfare economics and the theory of social discounting; and the evaluation methods literature, which is where the counterfactual designs this brief keeps asking for already live.
9 · Institutional requirements
Established Read the source list by interest before reading any number in it, because the pattern is unusually clean. Almost every positive claim about planning originates with a planning body, a treaty secretariat, a repository operator or a government describing its own machinery. Almost every adverse finding originates with an audit office, a parliamentary committee, a statutory inquiry or a peer-reviewed journal. Two sources here are load-bearing precisely because their interest runs against their finding: the study commissioned by a national foresight body which reported that the impact evidence base is limited, and the accident investigation run by the department whose oversight it faults.
Established The base rate explains most of the thinness, and it is not specific to long horizons. The UK's National Audit Office found that in 2019 only 8% of government spend on major projects — £35bn of £432bn — had robust evaluation plans in place. Of the 108 most complex and strategically significant projects, just 9 were robustly evaluated, while 77, representing 64% of spend, had no evaluation arrangements at all. Government “does not hold data on how far business as usual activities are covered by evaluation”. Established If 8% of near-term, high-salience, well-funded projects are evaluated, the prior probability that a fifty-year plan has been evaluated is close to zero, and the observed emptiness of this literature needs no further explanation than that.
Established At the global level the newest instrument commits nobody to anything measurable, and it is weaker than its own summaries suggest. The Declaration on Future Generations adopted in September 2024 contains nine action items across paragraphs 24 to 32, none of them binding, measurable or time-bound. It does not create a Special Envoy for Future Generations: paragraph 32(a) “takes note of” the proposal. It provides for a high-level review during the General Assembly's 83rd session with a Secretary-General implementation report, so the first structured occasion on which anything could be measured is around 2028–29. Frontier The correction matters in the direction that weakens the instrument, which is the direction that strengthens this brief's reading of it.
Frontier The institutional population itself is small, fragile and mostly undocumented. Sixteen extant proxy bodies across eleven countries; one abolished after a single term on operating-cost grounds; one surviving by demotion to a deputy within a general human-rights office; one long-running parliamentary committee whose own government-linked account describes what it does and presents no evidence that its recommendations changed legislation. None of these has an impact evaluation attached. Long-Term Institutions carries the mortality record in detail; what belongs here is the consequence for measurement, which is that the sample of institutions is 25 and the sample of evaluations is 1.
10 · Ethical & societal considerations
Established Every civilizational plan contains a rule for trading present against future, there is no agreed rule, and the disagreement is large enough to invert policy conclusions. The most-cited dispute sets the Stern Review against the DICE model. Stern's 0.1% is a pure rate of time preference, not a social discount rate, and the distinction is the whole argument: with an elasticity of marginal utility of 1, the corresponding consumption discount rate is about 1.4%. Nordhaus's model used 3% declining to about 1% over 300 years, calibrated to observed market rates, savings behaviour and capital returns. The consequence is an order of magnitude: an optimal 2005 carbon price of $17.12 per ton of carbon under the calibrated model against $159 under Stern's parameters, and a shift from a gradual policy ramp to 50% cuts by 2015.
Established Nordhaus's reductio is the sharpest single statement of what is at stake in the parameter. Under Stern's framework, a damage of 0.01% of output beginning in 2200 and continuing indefinitely would justify spending about 15% of today's global consumption — roughly $7 trillion — to prevent it. His conclusion is that the Review's “unambiguous conclusions about the need for extreme immediate action will not survive the substitution of discounting assumptions that are consistent with today's market place”. Frontier He is a party to the dispute and the reductio is an argument rather than a measurement; it is carried because it states the mechanism exactly.
Established The expert distribution is the closest thing to a settled position, and it is settled in the aggregate and unsettled in the parameters. Drupp, Freeman, Groom and Nesje surveyed economists who had published on social discounting since 2000: 262 responses from 627 identified, 197 experts including 12 who answered qualitatively only, with the parameter distributions resting on N of roughly 181 to 185. Recommended social discount rate: mean 2.27%, median and mode 2%, range 0% to 10%. Pure rate of time preference: mean 1.10%, median 0.50%, mode 0%, range 0% to 8%. Elasticity of marginal utility: mean 1.35, median and mode 1.0, range 0 to 5. Established There is real convergence — 92% found 1–3% acceptable and 77% were comfortable with 2% — and Stern's implied consumption rate of about 1.4% sits inside that acceptable band rather than outside it. Setting 0.1% against this distribution understates Stern by an order of magnitude and is the most common error in the popularisation of this dispute.
Frontier The ethical problem underneath is that the people a civilizational plan is for cannot be consulted, and the two available substitutes both fail differently. A proxy institution can be defunded, demoted or abolished by the generation it constrains — the mortality record says so. A discount parameter converts the question into arithmetic and hides the ethical choice inside a technical guidance document. Speculative The deliberative device tested in Japan is the only third option with any experimental support, and it has been tested on choices rather than on institutions. That is the honest state of the ethics: three mechanisms, one measured, and the measured one measured at the wrong scale.
11 · Civilizational implications
Established The most important result in this subject is that accurate long-range projection and effective long-range response are separable capabilities, and we demonstrably have some of the first. The 1972 Limits to Growth World3 scenarios have now been scored against outturn four times, not twice. Turner (2008) at CSIRO compared 1970–2000 UN, FAO, UNESCO and WorldWatch data using normalised RMSD and found the observed data “most closely match the simulated results of the LtG standard run scenario for almost all the outputs reported” — year-2000 differences of 0% for population, 5% for industrial output per capita, 5% for food per capita, 15% for pollution and 5–25% for non-renewable resources, with the crude death rate the notable miss at 40%. Turner published an update in 2014. Herrington (2021) extended the comparison to 2019 across ten variables and two accuracy measures and found the closest fits were the BAU2 and CT scenarios. Nebel, Kling, Willamowski and Schell added a fourth in the Journal of Industrial Ecology in 2023 — carrying a Correction published in May 2025, which anyone citing it must carry too.
Established Read that literature the way its authors ask and it is a finding about planning, not about doom. Turner is explicit that “all LtG scenarios show the global economic system growing at the year 2000”; Herrington is explicit that the comparison demonstrates World3's merit as an analysis tool for general global dynamics rather than for point prediction. Established And the response record is the inverse of the projection record. A 1972 projection tracked reality for decades and was read by every government in the world; over the same five decades global material extraction tripled, and the UN environment programme's resource panel projects it rising a further 60% by 2060 from 2020 levels, with rising trends in global resource use “continued or accelerated” since 2019.
Established So the terminal position of this brief is a split, and it is a real one rather than a failure to conclude. On projection: measured, replicated, partly successful, and improving slowly or not at all. On execution: one plan at civilizational horizon with independent verification and a modelled counterfactual, two more at sub-generational and generational horizon with measured outturns and real verification, everything else self-scored, and one institution externally examined with the finding that it was not driving the change intended. “Civilizations can plan at civilizational scale” is not refuted by this record. It is unmeasured, in a specific and fixable way.
Frontier What this page is for, stated plainly, is that the gap is narrow enough to close and nobody has closed it. The missing study takes a statutory long-horizon planning institution, is run by a body with no stake in the answer, and builds a counterfactual against outcome data rather than scoring process. The Wales-versus-rest-of-UK natural experiment is the obvious instance and it remains unrun at the institutional level. Frontier If the ESRC first-stage evaluation or the Senedd post-legislative inquiry has landed with outcome attribution, that plus the ozone treaty is two cases and this slot becomes arguable rather than ongoing. Established That is the condition on which this brief is revised, and it is written here so a reader can check it before this page's authors do.
Speculative The long-run implication, if the durability hypothesis turns out to be right, is that the binding constraint is coalition survival rather than analysis. On that reading a civilization's planning capability is a property of its political system's ability to keep a commitment alive across the people who made it, and modelling capacity, forecasting accuracy and engineering are all downstream. Handwave Nothing in the located record measures coalition survival — there is no denominator, no hazard model and no controlled comparison — so this remains the most consequential untested proposition in the subject, and it is named here rather than asserted.
12 · Timelines
These horizons track evaluation decisions, statutory review dates and one atmospheric recovery curve rather than technology:
- 10 yr: Frontier The two publications that would move this slot are already commissioned: a first-stage evaluation of the Welsh Act by an ESRC policy fellow, and a Senedd post-legislative inquiry. Established The Act's own auditor has asked the Welsh Government to set a scope and timetable for post-legislative evaluation; whether that happens is checkable and dated. Established At the global level the first structured occasion to measure the Declaration on Future Generations is a high-level review in the General Assembly's 83rd session, around 2028–29. Frontier And the Finnish repository either receives an operating licence or does not, which converts a forty-three-year programme from a plan into an operation.
- 25 yr: Frontier Ozone return to 1980 values at Northern mid-latitudes is projected for around 2035 and near-globally for around 2040, which is the first time a civilizational-scale plan will be verifiable against its own completion criterion rather than against a projection. Speculative The plausible institutional split is that plan attainment continues to be self-scored while the physical-verification cases accumulate, so the evidence base grows on the one problem shape that has an instrument and stays empty on the rest.
- 50 yr: Speculative Either an independent counterfactual evaluation of a long-horizon institution exists and this page is rewritten as an assessment, or the field completes another fifty years of legislating without measuring and the correct description becomes political rather than epistemic. Speculative A repository closes and markers are erected somewhere in this window, converting the message-preservation literature from design study into an untestable artefact. Handwave Which of those happens is not forecastable from anything in the current record.
- 100 / 250+ yr: Frontier Antarctic ozone return is projected for around 2066 — the only dated completion criterion in this brief that lies beyond a career, and the single most useful scheduled observation in the subject. Handwave Beyond it, nothing here forecasts: the longest institutional commitment with a documented outturn is forty-three years old and still pending, and the ten-thousand-year regulatory periods are regulatory constructions rather than measurements.
13 · Technology tree & dependencies
- Depends on Nothing on this map. No typed depends-on edge is claimed, and the absence is a finding rather than an omission: civilizational planning waits on no result another brief produces. Its blockers are measurement decisions that audit offices, legislatures and research councils could take tomorrow, and they are recorded as typed requirements instead.
- Requires (not on this map) Four institutional supplies, none of them a research result. First, an evaluator with no stake in the institution it evaluates: the one statutory future-generations duty examined at scale was examined by the audit office of the same country, which is the strongest arrangement anyone has built and is still not an outside party. Second, a counterfactual design applied to that duty — Wales legislated in 2015 and the rest of the United Kingdom did not, giving a control group and roughly eleven years of comparable indicator data, and at the institutional level nobody has run it; the Welsh Government's own statistical framework assesses direction of change only and explicitly declines to ask whether milestones are on course, so attribution has to be built rather than read off. Third, external verification of national plan attainment: attainment rates of 20 of 23, 28 of 28 and 24 of 33 are scored using data collected by the bodies being scored, with evaluation criteria negotiable at the outset on the testimony of a State Department analyst, and no audit office anywhere verifies a multi-decade national plan. Fourth, a survival denominator: twenty-five proxy institutions are known, sixteen survive, one was abolished after a single term and one was demoted, and there is no record of how many multi-decade statutory commitments have been made, how many survived a change of government, or what predicted survival — which is why the durability hypothesis in section 3 stays a hypothesis. Each of the four is something a legislature, an audit office or a research council could choose to supply, and each has so far not been supplied.
- Enables In principle, every brief on this map with a multi-decade horizon rests on the proposition that a plan can outlive its authors. No typed enabling edge is claimed, for the same reason a verdict is not: the enabling relationship is exactly the thing that has been evaluated once, without a counterfactual, with the finding that the intended system-wide change was not happening.
- Adjacent Atmospheric chemistry and the ozone assessment process, which supplies the only independent long-horizon measurement here; the forecasting and judgement literature; welfare economics and the theory of social discounting; evaluation methods, where the missing counterfactual designs already exist; and within this map Public Policy Foresight, Long-Term Institutions and Megaproject Governance.
14 · Common misconceptions & speculative claims
Established “Wales proves statutory future-generations duties work.” It is the most-cited case because it is the only one with an audit trail, and the audit trail says the opposite: after approximately 200 audit reports over five years the Auditor General concluded that “the Act is not driving the system-wide change that was intended”. Established And “Wales proves they don't work” fails on the same evidence. No counterfactual has been constructed, and the Welsh Government's statistical framework measures direction of change while explicitly declining attribution. “No measured effect” here means not measured, not measured as zero, and the difference is the reason this page carries the status it does.
Established “The CFC-11 episode shows the ozone treaty's monitoring failed.” The cited assessment records the opposite sequence: detection by the monitoring network, regional attribution, and a decline in global CFC-11 emissions after 2018 with eastern China explaining 60 ± 30% of the decrease. The delay it cost is bounded — about one year for mid-latitude equivalent effective stratospheric chlorine and up to three years at the poles — and quoting only the larger figure misreports the source. This is the one worked example in the record of a civilizational plan detecting and correcting a violation, and it should be cited as one.
Speculative “Finland succeeded and the United States failed because of politics, and the geology was comparable.” The first half is a live and interesting hypothesis; the second half is unsupported. Nothing cited here establishes that the two repository problems were technically comparable — unsaturated oxidising tuff under a court-imposed million-year standard and saturated crystalline bedrock with copper canisters are different problems, and no source in this brief compares them. Handwave Calling a two-case comparison “the measured variable” is the step where the argument does the work by assertion, and a brief that declines to generalise from one ozone treaty cannot licence it.
Established “Nobody has ever tested whether representing future generations changes decisions.” False, and the correction is more interesting than the claim. Japan's Future Design programme has run controlled experiments: 60% chose the sustainable option facing an imaginary future generation against 28% without. Frontier The accurate statement is that the mechanism has been tested at the deliberative level and not at the institutional level — which is why this brief's status turns on institutions rather than on whether the underlying device does anything.
Established “Government foresight has never been evaluated.” Also false. There is an established foresight-evaluation sub-field — Georghiou and Keenan (2006); Georghiou, Keenan and Miles (2010) on the UK programme itself; Havas, Schartinger and Weber (2010); Dursun, Türe and Daim (2011); Seidl da Fonseca (2016). Established What is missing is a particular kind of evaluation, not evaluation as such: that literature scores process, network formation, participant self-reported use and behavioural additionality, and does not score decisions against a counterfactual. Anyone repeating a count of the form “N reviews, zero found impact” is making a falsifiable claim that a literature search overturns.
Established “China's five-year plans prove that planning capability is real.” The attainment figures are real and the inference is not: the data used to score performance is collected by the bodies being scored, evaluation criteria are negotiable at the outset, and no independent body verifies attainment. Established “The Limits to Growth retrospectives prove planning works” fails in the opposite direction — they are evidence that projection worked while response did not, since global material extraction tripled across the same five decades in which the projection tracked outturn.
Established “The Declaration on Future Generations created a Special Envoy.” It did not. Paragraph 32(a) takes note of the proposal; the Declaration's nine action items contain no targets, deadlines, enforcement mechanisms or compliance criteria. Established “Repositories demonstrate institutional memory across ten thousand years” is stronger still and equally unfounded: no repository has been closed, no marker system has been built, and the record shows a facility licensed for ten-thousand-year isolation taken offline for three years, at year fifteen, over a repackaging procedure written at another site.
Established “Stern used a 0.1% social discount rate.” He used a 0.1% pure rate of time preference; with an elasticity of marginal utility of 1 the implied consumption discount rate is about 1.4%, which sits inside the band 92% of surveyed experts found acceptable. Setting 0.1% against a distribution of social discount rates understates his position by an order of magnitude. Established And the expert survey's parameter distributions rest on roughly 181 to 185 respondents, not the 197 experts or the 262 responses the paper also reports — three different Ns for three different things, routinely collapsed into one.
Frontier “Tsunami stones and disaster memorials prove multi-century warning transmission works.” The available peer-reviewed measurement surveys people who chose to visit a memorial, which is selection on the outcome; the popular claim that stone markers saved villages in 2011 circulates in press accounts and could not be substantiated here from any source establishing a counterfactual. Established And none of the emptiness in this subject should be read as a fact about long horizons specifically: an audit office found that 9 of a country's 108 most significant projects were robustly evaluated and 77, representing 64% of spend, had no evaluation arrangements at all. The absence of evaluation of civilizational plans is a special case of a near-universal absence of evaluation.