1 · Concept overview
This brief is about composite country-level indices — the Human Development Index, the global Multidimensional Poverty Index, the Sustainable Development Goal indicator framework, and the wider class of global performance rankings — and about what publishing them does to policy. It sits with two others on this map where the same structure recurs: a measurement has come to stand in for the thing it names, and has acquired institutional force the thing never had. In Human Flourishing the thing is wellbeing; in Reputation Economies it is trustworthiness; here it is development.
The framing under test — that development can be captured in an index that guides decisions — again contains two claims. (a) The composite captures development, in the sense that it is a defensible summary of the thing. (b) The index guides decisions, in the sense that publishing it changes what governments do in a direction someone can measure. Frontier The surprise of this brief, and its centre, is that the second claim is better supported than the first. Indices demonstrably move states, sometimes at head-of-government level. What they move states toward is a separate question, and the best-documented answer is a ranking that was gamed, manipulated and then discontinued by the institution that published it.
Established One boundary has to be stated at the top because it is routinely got wrong. The ordinality objection that dominates the wellbeing brief next door — that reported happiness is a categorical response whose group means are not invariant to cardinalisation — does not transfer to this subject. Life expectancy, years of schooling, gross national income per capita, whether a household cooks on solid fuel: these are objective quantities on real scales. The HDI and the MPI have a weighting problem, not a cardinality problem, and it is a different failure with different remedies. Confusing the two makes both critiques weaker than they are.
2 · Current scientific position
Established State the HDI exactly, because almost every argument about it is conducted without its arithmetic. From the publisher's own 2025 technical notes: health enters as life expectancy at birth normalised between goalposts of 20 and 85 years; education enters as expected years of schooling between 0 and 18 and mean years of schooling between 0 and 15, combined by arithmetic mean; standard of living enters as GNI per capita in 2021 PPP dollars between $100 and $75,000, in logarithms, on the stated ground that the transformation from income to capabilities is likely to be concave. The three dimension indices are then combined by geometric mean. The income maximum is set at $75,000 explicitly so that additional income above the maximum does not contribute to a country's value or improve its ranking.
Established Three of those decisions carry most of the weight and are almost never noticed by users of the index. First, the goalposts are chosen and consequential: a life-expectancy ceiling of 85 and a schooling ceiling of 18 years fix the whole scale, and every country's score is a position between two numbers somebody picked. Second, the geometric mean was adopted in 2010 specifically to reduce substitutability between dimensions — a deliberate change to the implicit tradeoffs. Third, the $75,000 cap combined with logarithmic income means the index is by construction insensitive to most of the variation in world income. Established UNDP states the limits itself, in plain terms: the HDI “simplifies and captures only part of what human development entails” and “does not reflect on inequalities, poverty, human security, empowerment, etc.” Its technical notes say the same of the satellite indices — the Inequality-adjusted HDI “is not association sensitive, so it does not capture overlapping inequalities”, the Gender Inequality Index likewise, and both cover fewer countries than the headline index. The publisher documenting its own instrument's limits runs against interest and is weighted up accordingly.
Established The sharpest technical objection is Ravallion's, and it is not answered. His argument is that the 2010 revision, in relaxing perfect substitutability, created new and troubling implicit tradeoffs: “a poor country experiencing falling life expectancy due to (say) a collapse in its health-care system could still see its HDI improve with even a small rate of economic growth.” The revised index therefore assigns lower implicit weight to mortality in poor countries than in rich ones, and he adds that the new index's valuations of extra schooling look “unreasonably high — many times greater than the economic returns to schooling”. Frontier Only the abstract was retrievable for this brief, so the specific monetary valuations he derives are deliberately not quoted here. Established Stated plainly, the finding is this: an index built to assert that people matter more than GDP contains, inside its aggregation rule, a lower price on a poor person's life than on a rich person's — and that price was set by the index designer, not by anyone accountable.
Established The Multidimensional Poverty Index is the strongest counterexample to blanket index scepticism, and its problems are of a different kind. Built on the Alkire-Foster method: three equally weighted dimensions over 10 indicators — health (nutrition, child mortality), education (years of schooling, school attendance), and living standards (cooking fuel, sanitation, drinking water, electricity, housing, assets) — equally weighted within dimensions, so each health and education indicator carries 1/6 and each living-standards indicator 1/18. A person is multidimensionally poor if deprived in at least one-third of weighted indicators. The 2025 round covers 109 countries and about 6.3 billion people, of whom 18.3 per cent are multidimensionally poor, with data updated for 13 countries or territories covering about 6.9 per cent of the developing world's population. Established What it does better than the HDI is structural: it is built up from individual-level joint deprivations rather than country averages, so it knows whether the same person is deprived in health and in schooling — the association sensitivity UNDP's own notes concede the IHDI and GII lack.
Established And its publisher documents its currency problem against interest. The surveys underlying the 2025 index span 2013–14 through 2024. Sixty countries, hosting 62 per cent of the multidimensionally poor, have data from 2019–20 onward; the rest is built on surveys up to a decade old, and cross-time comparison requires separate harmonisation protocols. Established What it shares with the HDI is the unvoted judgement: the weights are chosen, the one-third cutoff is chosen, and those two choices between them determine who counts as poor. Nobody voted on 33.33 per cent.
Established The SDG framework's tier classification is the cleanest case of a statistic meaning less than it sounds. As of 10 April 2025 the global framework comprises 234 indicators: 161 at Tier I (68.8 per cent), 60 at Tier II (25.6 per cent), 8 at multiple tiers and 5 pending data availability. Established Read the Tier I definition, because the definition is the finding. Tier I means the indicator is conceptually clear, has an internationally established methodology and available standards, and data are regularly produced by countries for at least 50 per cent of countries and of the population in every region. So the best tier — 69 per cent of indicators — is defined by a 50 per cent country-coverage floor, and an indicator can be Tier I with half the world's countries reporting nothing. Tier II has an agreed methodology and no regular country data production at all. Established To be fair to the framework: Tier III, indicators with no agreed methodology, has been worked down to zero, and the 2025 review reclassified four indicators upward. The methodological development succeeded. The data production did not keep pace with it.
Established Underneath every index sits national statistical capacity, and it is weaker than the indices' decimal places imply. Nigeria's April 2014 GDP rebasing found the economy nearly 90 per cent larger than previously measured; Ghana's 2010 rebasing produced a 60 per cent upward adjustment that moved the country into middle-income status, with consequences for its access to development finance. Frontier The development think tank reporting it draws the obvious implication — that thousands of econometric studies of Sub-Saharan Africa relied on the pre-revision data — and it is an interested party with a standing case for better statistics, so the interest runs with the claim. Established The structural point does not depend on the interpretation: a GNI per capita input that can move 90 per cent in a single administrative revision is one of the HDI's three components, entering in logarithms against a $100 to $75,000 goalpost range. The apparent precision of an index is not a property of what is underneath it.
Established Now the centre of this brief, because it is the one place the full causal chain is documented and most of the documentation comes from the publisher. Doing Business ran the sequence index published, governments respond, index distorted, index destroyed, and every link is on the record. Governments responded at the highest level: the World Bank's own Independent Evaluation Group records that Indonesia's president set a target of reaching 40th place, Russia's leadership ordered 50th place by 2015, and Rwanda continually pursued “top reformer” status, with IEG assessing these as ranking-competition targets “not necessarily outcomes benefiting small businesses”. Frontier Independent work reports more than 70 countries forming regulatory reform committees oriented to the indicators, and documents Kosovo receiving USAID-funded consultancy from 2010 to 2018 with a top-40 position as “a core objective” — achieved in Doing Business 2018 — alongside more than $100 million of comparable business-environment funding across Azerbaijan, Georgia, Kazakhstan, Kyrgyzstan, Moldova, Serbia and Ukraine.
Established The index did not measure what it claimed, and the publisher's own evaluators said so. IEG found the correlation between Doing Business's construction-permit times and actual firm experience in surveys to be 23 per cent, and for electricity connections 10 per cent, with the indicators tending to overestimate time compared with actual firm experience. Among World Bank practitioners interviewed, only 60 per cent thought the indicators validly measured their business area and only 26 per cent considered Doing Business a good way to monitor or evaluate reforms. IEG also notes that the binding constraints in the model reformer, Rwanda — low skills, limited energy access, high transport costs, restricted land access — were mostly outside the index's scope. Frontier And the rankings moved in ways that look artefactual: over fifteen years 111 countries experienced ranking spreads exceeding 40 positions and 35 countries fluctuated by 75 or more, while the methodology changed repeatedly, with nine of ten indicator categories changing at the 2015 shift to the distance-to-frontier method.
Established Then the data were manipulated, and the timeline is the publisher's own. Data irregularities were first reported internally in June 2020; the World Bank paused publication on 27 August 2020 and discontinued the report on 16 September 2021 after audits and an independent external review, with its President stating the Bank would develop replacements that are “trustworthy” and “confidence-inspiring”. The irregularities affected reports from 2016 to 2020. Frontier The countries named in the reporting are China (Doing Business 2018) and Saudi Arabia, the United Arab Emirates and Azerbaijan (2020); for China the alleged methods included combining Hong Kong, Taiwan and Macao scores into China's index, using only the better-performing of Beijing or Shanghai rather than averaging, and discarding less favourable expert assessments. Those country-level details rest on an academic's newspaper analysis of the investigation rather than on the investigation report, which could not be retrieved, and are flagged frontier for that reason. The successor product, Business Ready, launched in 2024.
Established Four things are documented about that episode and they should be stated together, because each is often quoted without the others: the index produced genuine reform effort; the effort was aimed at the indicator rather than the outcome; the measurement correlated weakly with the reality it named; and the publisher's own data were manipulated. This is the strongest single Goodhart case in the development-indicator record and it is stronger than anything available for the HDI or the MPI.
Frontier The best general evidence that indices change policy comes from outside development indices entirely. Kelley and Simmons, in the American Journal of Political Science, study the US State Department's Trafficking in Persons ratings and find that inclusion in the report is associated with increased criminalisation of human trafficking by rated countries, with watch-list placement likewise prompting criminalisation; their theory is that numerical ranking and public comparison constitute “an exercise of social power” operating through reputational incentives rather than enforcement. Verified from the abstract only, so no effect size is quoted. Established Alongside it, the clearest institutional adoption of a development index: Colombia launched a national MPI in August 2011, formalised it through CONPES 150 in 2012 with DANE as official producer, uses 15 weighted indicators with a one-third cutoff, feeds it into the National Development Plan, publishes annually, extended it to a Municipal MPI in 2020, and has chaired the Multidimensional Poverty Peer Network steering committee since 2013. Reported by the method's co-developer, so an interested party. Frontier And the honest summary: rankings change state behaviour, a development index can be institutionalised into a national planning system, and no source retrieved for this brief documents a specific policy decision changed by the HDI or the MPI with the counterfactual on record.
3 · Frontier questions
Frontier Question one: is a composite a defensible summary of development at all, and by what criterion? The honest answer is that there is no agreed criterion for what a good composite must preserve, so the question cannot currently be adjudicated. UNDP's own position is weak as stated — the index captures only part of what human development entails and does not reflect inequality, poverty, human security or empowerment. Established What is settled is that composites embed unchosen value judgements: HDI goalposts of 20–85 years and $100–$75,000, a geometric mean, logarithmic income, the MPI's 33.33 per cent cutoff and 1/18 living-standards weights. The dispute is not whether the judgements are there. It is whether they matter.
Frontier Question two: is the MPI a better composite than the HDI, and can that be tested? The structural case is real — association sensitivity at the individual level, which UNDP concedes its own adjusted indices lack. Frontier The test nobody has run is whether MPI rankings are more stable across editions and more predictive of later outcomes than HDI rankings. Both indices are published annually with long back-runs; this is a computation, not a research programme.
Frontier Question three: does the Doing Business Goodhart result generalise? The strongest argument that it does not is specific and should be given its weight: Doing Business was uniquely salient, uniquely tied to capital flows, and uniquely gameable through binary regulatory checkboxes, while the HDI's components — life expectancy, schooling, income — are much harder to game and much slower to move. Speculative The strongest argument that it does is that salience is a policy variable rather than a fixed property, and that any index which acquires Doing Business's salience acquires its incentives. Nobody has measured gaming in a slow-moving composite, so the generalisation is untested in both directions.
Frontier Question four, the minority position with real institutional traction: should dashboards replace composites? Report the components and refuse to aggregate, on the ground that aggregation is exactly where the unvoted judgements enter. The objection is empirical rather than philosophical: dashboards do not generate the salience that produces the behaviour change in the first place. A ranking creates a rank; a dashboard creates a table. Speculative Question five, further out: non-compensatory aggregation, in which no dimension may substitute for another, which would dissolve Ravallion's tradeoff problem by construction. Proposals exist in the methodological literature; adoption is nil. Publishing a parallel non-compensatory HDI and comparing rankings would cost one analyst a month.
Frontier Question six, the fringe position that is harder to dismiss than it looks: are rankings instruments of foreign policy? The observation behind it is non-trivial — the two best-documented cases of rankings changing state behaviour are a US government rating of other states and a World Bank product whose manipulation is alleged to have favoured China and Gulf states. Speculative That is suggestive and not demonstrated; what would settle it is a systematic analysis of whose scores improve when publishers face political pressure, and no such analysis exists. Speculative Question seven, from the same neighbourhood: is index-driven reform net harmful because it diverts scarce administrative capacity from binding constraints to indicator management? IEG's Rwanda finding implies it — the binding constraints were outside the index's scope while indicator reform absorbed sustained government attention — and nobody has measured administrative effort against outcomes.
4 · Technological bottlenecks
Established The first bottleneck is that the inputs move more than the outputs are supposed to discriminate. An index reports to three decimal places on a component that can be revised by 60 or 90 per cent. Nothing about the aggregation can repair that, and no published index attaches an uncertainty interval to a country's rank. Frontier Rank is the quantity everyone consumes and the quantity nobody publishes an error bar for.
Established The second is coverage masquerading as capability. The SDG framework's headline tier statistic is a statement about methodology plus a 50 per cent country floor, not a statement about measurability. What is missing is trivial to produce and has not been produced: actual per-indicator country-coverage rates, published alongside the tier counts. Established The MPI's version of the same problem is vintage: a global index in 2025 built partly on surveys from 2013–14, with 62 per cent of the poor covered by data from 2019–20 onward and the remainder older.
Established The third is that the aggregation function is a value judgement with no accountable author. Goalposts, means, caps, cutoffs and weights are all chosen; the choices determine rankings and, in the MPI's case, who is counted poor at all. Frontier There is no procedure anywhere in the record by which an affected population consents to a weighting. That is a governance gap, not a statistical one, and it is why the critique keeps having to be made from outside.
Frontier The fourth is comparability across editions. Doing Business changed nine of ten indicator categories in a single methodological shift while 111 countries moved more than 40 ranking positions over fifteen years. Every composite revises its method; almost none republishes the full back-series on the new method. Established A rank change that is a method change looks identical, in a headline, to a rank change that is a country change.
5 · Research dependencies
Established Nothing on this map produces a result this brief waits on. The constraints are institutional and are recorded as typed requirements below: national statistical capacity that does not move by tens of per cent on revision, per-indicator coverage published rather than tier counts, and an index publisher insulated from the states it ranks. None of the three is a discovery. All three are choices somebody could make.
Frontier What this brief waits on from research is narrow and testable. A comparison of ranking stability and predictive validity between the MPI and the HDI; a systematic comparison of index rankings before and after major national-accounts rebasings; and a measurement of whether gaming appears in slow-moving composites the way it appeared in a fast-moving one. Speculative All three use data that already exists.
Speculative Two areas are empty rather than thin. There is no study documenting a specific allocation decision changed by the HDI or the MPI with the counterfactual on record — institutional adoption, including Colombia's CONPES 150 and its Municipal MPI, is real and is not the same thing. And there is no analysis of whose rankings improve when index publishers come under political pressure, which is what the foreign-policy critique in section 3 would need to become more than an observation.
6 · Required experiments
Established The cheapest high-value study is a recomputation, not a survey. Publish a parallel HDI under alternative aggregation — non-compensatory, different goalposts, no income cap — and report how far rankings move. Frontier If the ordering is robust to the aggregation rule, the strongest critique of composites loses most of its force; if it is not, every published ranking carries an unstated dependence on choices its users never see. Either result is worth having and neither requires new data.
Established Second: put an error bar on a rank. Take the rebasing record — Nigeria's 90 per cent, Ghana's 60 per cent — recompute the affected countries' HDI positions before and after, and publish the induced rank movement as an estimate of input uncertainty. Frontier This has been available to do since 2014 and would convert a rhetorical objection into a number.
Established Third: publish per-indicator country coverage for the SDG framework. The tier classification already encodes the information; reporting the actual share of countries producing data for each indicator, rather than the count of indicators clearing a 50 per cent floor, would replace a misleading headline with an informative one at essentially zero cost. Frontier Fourth: run the MPI-against-HDI horse race on ranking stability across editions and on predictive validity for later outcomes, using the published back-series of both.
Frontier Fifth, the one design that would settle the framing's second clause: compare budget allocations before and after national MPI adoption against a credible comparison group. Colombia adopted in 2011–12 and extended to municipalities in 2020, and a number of other countries have adopted national MPIs on staggered dates. Speculative That is a difference-in-differences with real treatment timing, and nobody has run it. Sixth: register in advance what gaming would look like for a slow composite — enrolment inflation, selective survey timing, boundary redefinition — before an index acquires the salience that would produce it.
7 · Engineering requirements
Established The engineering here is the aggregation rule, and it is more consequential than any data-collection decision downstream of it. The HDI's pipeline is four indicators, four pairs of goalposts, one arithmetic mean, one logarithm, one cap and one geometric mean. Change the geometric mean back to an arithmetic one and dimensions become perfectly substitutable again. Lift the $75,000 cap and rich-country ordering changes. Move the life-expectancy ceiling from 85 and every country's health index moves. Frontier None of those parameters is derived from anything; each is a defensible choice, and the index reports the result to three decimals.
Established The MPI's pipeline is a dual-cutoff counting method and the second cutoff does the work. Ten indicators, weights of 1/6 and 1/18, and a poverty cutoff at one-third of weighted deprivations. The first cutoff decides whether a household is deprived on an indicator; the second decides whether enough deprivations make a person poor. Established Its genuine engineering advantage is that the counting happens at the person, so joint deprivations are visible — the property UNDP's technical notes state its own adjusted indices lack. Its unremovable judgement is the same 33.33 per cent that decides the headline.
Established The SDG framework is an indicator registry rather than a composite, and its engineering problem is production rather than aggregation. 234 indicators; Tier III worked down to zero, which is a real methodological achievement; and a Tier I definition that requires regular country data production for only half of countries and population in every region. Frontier A registry that grades its indicators on methodology and coverage, and then reports the grade rather than the coverage, has built an instrument that reads as more complete than it is.
Established And the input layer is the weakest engineering in the stack. National accounts revisions of 60 and 90 per cent; household surveys up to a decade old inside a current index; per-indicator coverage unpublished. Frontier Everything above this layer is arithmetic on numbers whose uncertainty nobody propagates. No composite in the record carries an uncertainty interval from input revision through to published rank, and constructing one is a tractable engineering task that has not been done.
8 · Adjacent technologies
The boundary that matters most on this map is with the wellbeing brief, and it is resolved by unit of analysis rather than by subject. Human Flourishing owns person-level wellbeing and its domains: life-satisfaction scale validity, the Easterlin dispute, the WELLBY, and Bhutan's Gross National Happiness, which is assigned there because its arithmetic runs over persons. This brief owns country-level composite indices and what publishing them does to policy, and may name GNH as a composite without restating its survey results. The Alkire-Foster method appears in both because both objects use it; here it is cited for the MPI. Established And the ordinality objection that dominates next door does not transfer here. Life expectancy, schooling years and GNI per capita are objective quantities on real scales, so the HDI's and MPI's difficulty is a weighting problem and not a cardinality one — a different failure with different remedies, and worth stating explicitly because the two critiques are constantly merged.
Within this map, also: Reputation Economies, the third of the three measurement briefs, where the Goodhart pathology takes the form of inflation and counterfeiting rather than rank-chasing; Intelligence Measurement, the same aggregation question asked of a latent individual trait; Global Cooperation Models, which owns the institutions that publish and consume these rankings; Future Public Administration, whose evaluation-is-discretionary finding explains why index adoption is documented and index effects are not; Institutional Design, where the incentive problem measured here as rank-chasing is stated generally; Public Policy Foresight and Long-Term Institutions, for the horizon over which an indicator has to stay comparable; and Civilizational Planning, which inherits the question of what a civilisation should measure about itself.
Outside it: official statistics and national accounting, which supplies the inputs and the rebasing record; the global-performance-indicator literature in international relations, which supplies the evidence that rankings move states; development economics, which supplies the composites; and the sociology of quantification, which supplies the argument that measuring a thing changes it.
9 · Institutional requirements
Established The most striking institutional fact in this subject is how much of the damning evidence comes from the publishers. UNDP states that its own index does not reflect inequality, poverty, human security or empowerment, and that its adjusted indices are not association sensitive. OPHI states its own index's data-vintage problem. The World Bank's Independent Evaluation Group found its own flagship correlating 23 per cent and 10 per cent with surveyed firm experience, and that only 26 per cent of the Bank's own practitioners thought it a good way to monitor reform. The Bank then paused and discontinued the product and published the timeline. Established All of that runs against interest, which raises its weight substantially, and it is the behaviour of a serious statistical enterprise rather than a propaganda one.
Established The countervailing institutional fact is that the same publisher's data were manipulated for five consecutive editions. Irregularities affected reports from 2016 to 2020; publication was paused in August 2020 and the report discontinued in September 2021. Frontier An institution that both publishes a global ranking and lends to the states it ranks holds two positions that cannot be reconciled by disclosure, and the naming of specific beneficiaries rests on secondary reporting rather than the investigation report, which could not be retrieved for this brief.
Frontier The requirement that follows is structural rather than procedural: an index publisher insulated from the states it ranks. Insulation is not the same as independence in name — it means no lending relationship, no consultancy revenue from ranking improvement, and no ability for a rated state to reach the compilers. Established The counterexample in the record is instructive in both directions: donors funded consultancies whose “core objective” was reaching a top-40 rank, in a system where more than 70 countries had reform committees oriented to the same indicators. That is an ecosystem in which measuring and improving the measurement were performed by overlapping institutions.
Established Two further institutional requirements are recorded below because they are cheap and undone. Publishing per-indicator country coverage would replace a headline that means less than it sounds with one that means what it says. And statistical capacity that does not move by tens of per cent on revision is a funding question in national statistical offices, not a methodological one. Neither requires a research advance and neither has been supplied.
10 · Ethical & societal considerations
Established The sharpest ethical finding in this brief is Ravallion's, and it deserves stating without softening. An index constructed to insist that human lives matter more than national income embeds, in its aggregation rule, a lower implicit weight on mortality in poor countries than in rich ones. A poor country whose health system collapses can still record an improving score on modest growth. Frontier That is not a rhetorical charge; it is a property of the geometric mean applied across normalised dimensions with those goalposts, and the numerical valuations behind it are in a paper this pass could not retrieve, which is why the argument is carried and the figures are not.
Established The second is that the poverty cutoff decides who is poor. One-third of weighted deprivations is a defensible line and it is a line somebody drew; move it and the count of the multidimensionally poor moves with it, across 6.3 billion people. Frontier The people the number describes have no procedure for contesting the weights, and the index that determines their classification is published in a different country from almost all of them.
Frontier The third is what index competition does to administrations that can least afford it. IEG identified the binding constraints in the model reformer as skills, energy, transport and land access — mostly outside the index's scope — while the reform effort oriented to the index absorbed sustained attention at the top of government. Speculative If scarce administrative capacity is the real constraint in a poor state, then an index that captures that capacity has a cost that no evaluation of the reforms themselves will show. That is a hypothesis, and the fact that it is untested after two decades of ranking is itself a finding.
Established And the fourth is the manipulation itself. A ranking that shapes capital flows was altered for five editions, with the alleged beneficiaries being large and geopolitically weighty states. Frontier Whatever the intentions, the structural lesson is that the incentive to manipulate a global index falls hardest on the publisher, not on the ranked — and the ranked have no remedy, because they cannot audit the compilation.
11 · Civilizational implications
Established The terminal position is a tie and both halves must be held. Composite indices encode unaccountable value judgements on unreliable data: goalposts and caps nobody voted on, a cutoff that decides who is poor, inputs that move 60 and 90 per cent on revision, a best tier defined by a 50 per cent coverage floor, and an aggregation rule that prices a poor person's life below a rich person's. And they are among the very few instruments demonstrably capable of moving state behaviour: heads of government setting rank targets, more than seventy reform committees, criminalisation following a rating. Frontier The site should hold both, and note that holding both is exactly why they are dangerous. An instrument that does not work and does not move anything is harmless.
Established The change of question this brief lands on is the inversion in its framing. The usual worry about indices is that they are inaccurate. The measured record says the accuracy problem and the influence problem are independent, and the second is larger. Doing Business correlated 23 per cent with the reality it named and simultaneously reorganised regulatory policy in dozens of states. Frontier The gap between how well an index measures and how hard it pulls is where the damage happens, and nothing in the design of any current index closes it.
Frontier The long-run risk is not that a bad index misleads a decision but that a good index becomes the terrain. Once ranks are set as head-of-government targets, the index stops being a description of a country and becomes a specification for one. Speculative That is a claim about how measurement interacts with political attention rather than about statistics, and the one documented instance ended with the publisher cancelling the product. Whether that ending is the exception or the pattern is genuinely open, and this brief does not pretend to know.
Speculative And the largest open question is whether anything can occupy the position without inheriting the problem. Dashboards avoid the aggregation judgement and forfeit the salience that made indices consequential. Non-compensatory aggregation fixes the tradeoff defect and has never been adopted. Handwave A recommendation that reduces to “measure development better” is not actionable, and the evidence assembled here supports a narrower and more useful one: publish the uncertainty, publish the coverage, publish the alternative aggregations, and separate the compiler from the lender. Every one of those is available now.
12 · Timelines
These horizons track publication cycles, survey vintages and the replacement of a cancelled product rather than technology:
- 10 yr: Established The HDI and the global MPI continue annually and their methods are stable, so within this window the interesting movement is in the inputs rather than the indices — further national-accounts rebasings and the refreshing of the MPI's oldest surveys, currently reaching back to 2013–14. Frontier Business Ready, launched in 2024, is the live test of whether a successor product can carry the influence of the cancelled one without the incentives; expect its early editions to be watched for exactly the gaming the predecessor produced. Frontier The SDG framework reaches its 2030 terminal date inside this window, and the tier classification will be quoted as an account of what was measurable. Speculative Expect at least one national MPI adoption study to be attempted, because the treatment timing is staggered and the data exist.
- 25 yr: Speculative The plausible split is that composite indices persist because their salience is the only thing that reliably moves states, while the methodological literature converges on dashboards that governments continue to ignore. Speculative If any current index acquires Doing Business's salience — tied to capital flows and set as a head-of-government target — expect the same sequence, because nothing in the design of the current instruments prevents it. Handwave Which index that would be is a guess about political attention, not an extrapolation.
- 50 yr: Speculative If national statistical capacity improves enough that revisions stop moving components by tens of per cent, the strongest technical objection to composites weakens substantially and the remaining objection is purely about the aggregation judgement, which is normative and will not be settled by data. Speculative If it does not improve, the indices will still be published to three decimal places. Handwave Both are extrapolations from a funding trajectory nobody has forecast.
- 100 / 250+ yr: Handwave Beyond useful forecasting. The only datum at that horizon is that national income accounting is roughly ninety years old and was contested as a summary of national condition from its first decade, and that the composites in this brief were built as replacements for it within fifty years. Handwave One succession, in one measurement tradition, is a story rather than a base rate.
13 · Technology tree & dependencies
- Depends on Nothing on this map. No brief here produces a result this one waits on; the constraints are institutional choices about who compiles an index, what they publish alongside it, and how well the underlying national statistics are funded. No typed depends-on edge is claimed.
- Requires (not on this map) All four are institutional and none needs a discovery. An index publisher with no lending relationship to the states it ranks: the one fully documented case of a global index shaping policy ran through an institution that both ranked states and lent to them, its data irregularities affected the 2016 to 2020 editions, publication was paused on 27 August 2020 and the report discontinued on 16 September 2021, and the ecosystem around it included donor-funded consultancies whose core objective was a top-40 position. Per-indicator country coverage published alongside the SDG tier counts: Tier I requires data regularly produced for at least 50 per cent of countries and of the population in every region, so the widely quoted figure that 69 per cent of indicators are Tier I is read as a statement about measurability when it is a statement about a coverage floor; the underlying coverage is already known to the classifiers and is simply not reported. National statistical capacity that does not move a component by tens of per cent on revision: Nigeria's 2014 rebasing found the economy nearly 90 per cent larger and Ghana's 2010 rebasing 60 per cent larger, and GNI per capita is one of the HDI's three components. And an uncertainty interval published on a country rank: rank is the quantity every consumer of these indices actually uses, 111 countries moved more than 40 positions in fifteen years of one ranking while its methodology changed repeatedly, and no composite in the record propagates input uncertainty through to the published position.
- Enables In principle, every development programme on this map that justifies itself by a target or a rank rests on the composite that defines it. No typed enabling edge is claimed, and the reason is a finding rather than an omission: no source retrieved documents a specific allocation decision changed by the HDI or the MPI with the counterfactual on record, so the enabling relationship has been asserted far more often than it has been shown.
- Adjacent Official statistics and national accounting, which supplies the inputs and the rebasing record; the global-performance-indicator literature in international relations, which supplies the evidence that public rankings move states; development economics, which supplies the composites and the counting methods; the sociology of quantification, which supplies the argument that measuring changes the measured; and within this map Human Flourishing, Reputation Economies and Global Cooperation Models.
14 · Common misconceptions & speculative claims
Established “Sixty-nine per cent of SDG indicators are measurable.” This is the most consequential misreading in the subject and it is settled by reading the definition. Tier I requires an internationally established methodology and data regularly produced for at least 50 per cent of countries and of the population in every region. An indicator can sit in the best tier with half the world's countries reporting nothing, and the 26 per cent at Tier II have an agreed methodology and no regular country data production at all. Established The tier statistic describes methodological readiness plus a coverage floor. It has never described measurability, and the UN's own classification document says so.
Frontier “Doing Business proves that all development indices are gamed.” It does not, and this brief resists rounding the strongest case in the record up into a law. Doing Business was uniquely salient, uniquely tied to capital flows, and uniquely gameable through binary regulatory checkboxes; the HDI's components are life expectancy, schooling and income, which are hard to game and slow to move. Established What Doing Business does prove is that the mechanism exists and can run to completion — rank targets set by heads of government, more than 70 reform committees, a 23 per cent correlation with firm experience, manipulation across five editions, cancellation. Speculative Whether a slow composite would show the same pattern under the same salience is untested, and asserting either answer is asserting something nobody has measured.
Established “The HDI is a fairer measure than GDP because it puts people first.” The intention is real and the arithmetic complicates it. Income enters logarithmically and is capped at $75,000, so the index is deliberately insensitive to most world income variation — that part works as advertised. Established But the post-2010 aggregation gives lower implicit weight to mortality in poor countries than in rich ones, so a poor country with a collapsing health system can post an improving score on small growth. An index built to put people before income contains a lower price on a poor person's life, and the price was set by the designer.
Frontier “Index rankings can be compared across editions.” Usually not, and the record is stark: over fifteen years of one ranking, 111 countries moved more than 40 positions and 35 moved 75 or more, while nine of ten indicator categories changed in a single methodological shift. Established A rank change caused by a method change is indistinguishable, in a headline, from a rank change caused by a country change, and almost no publisher republishes the full back-series on the new method.
Established “Bond and Lang's ordinality problem undermines development indices too.” It does not transfer, and the reason is worth being precise about because the two critiques are constantly merged. That objection concerns reported categorical responses whose group means are not invariant to cardinalisation. Life expectancy at birth, mean years of schooling and GNI per capita are objective quantities on real scales, and whether a household cooks on solid fuel is a fact rather than a rating. Established Composite indices have a weighting problem — unvoted goalposts, means, caps and cutoffs — and not a cardinality one. Different defect, different fix; the wellbeing critique belongs next door.
Frontier “Colombia proves an index can change national policy.” Colombia proves an index can be institutionalised, which is a real and useful fact: a national MPI from 2011, formalised through CONPES 150 in 2012 with an official producer, 15 weighted indicators, annual publication, a National Development Plan link and a municipal extension in 2020. Frontier Institutional adoption is not a documented change in an allocation, and no source retrieved for this brief compares budget allocations before and after national MPI adoption against a comparison group. The study is available to run and has not been run.
Speculative “Rankings are just instruments of the states that publish them.” This is a fringe position and it is not dismissible, which is why it is stated here rather than omitted. The observation behind it is that the two best-documented cases of rankings changing behaviour are a US government rating of other states and a World Bank product whose alleged manipulation favoured China and Gulf states. Speculative Two cases are an observation, not a demonstration; what would convert it is a systematic analysis of whose scores improve when publishers face political pressure, and no such analysis exists in anything fetched here.
Established Three sources are deliberately not relied on, and naming them stops their absence reading as an oversight. The WilmerHale investigation report into the Doing Business irregularities could not be retrieved, so the named countries and alleged methods are carried from an academic's newspaper analysis and flagged frontier. The full text of Ravallion's critique could not be retrieved, so his argument is carried and his numerical valuations are not. And Kelley and Simmons was verified from its abstract only, so no effect size is quoted for the finding that ratings raise criminalisation. Speculative Each gap is in a place where a number would have strengthened this brief, which is precisely why the number is absent rather than approximated.
Established And the framing itself: “captured in an index that guides decisions” joins two clauses that turn out to be independent. The guiding is real, forceful and documented at head-of-government level. The capturing is what failed. Frontier An index that guided decisions with great force while correlating 23 per cent with the reality it named is not a measurement failure and a policy success. It is a single mechanism working exactly as it was built to, in a direction nobody chose.