1 · Concept overview
“Grand challenge” is now a standard instrument of science policy. A convening body names a hard problem, attaches a purse or a deadline or merely a title, and expects effort to arrive. Governments, foundations and private sponsors have run hundreds of these. This brief asks the only question that matters about an instrument: does it work, and how would anyone know?
The framing under test is that framing a problem as a grand challenge mobilises effort effectively. The evidence separates cleanly along that sentence. Effort mobilisation is real and measurable. Effectiveness — additional innovation that would not otherwise have happened — is barely measured at all. The instrument's own evaluation literature says so in terms.
The ARPA model, the programme-manager mechanism and the preconditions for replicating DARPA belong elsewhere on this site and are not re-derived here. Public Policy Foresight owns the parallel finding about government futures machinery: a field that declined to define an outcome measure. This brief owns the challenge framing itself as an instrument — prizes, challenge lists, and whether any of them has a measured effect.
Three findings cut, in ascending order of damage. The founding case is not a case: there was no Longitude Prize as anyone imagines it. The purse is not the mechanism: in the only large-sample historical study, doubling the money bought 11% more entrants and one extra gold medal bought 53%. And it is not cheap: DARPA's own report to Congress records about $7.8 million of agency spending to award a $2 million prize.
2 · Current scientific position
Established Start with where the phrase comes from, because its genealogy is not what its users think. David Kaldewey, The Grand Challenges Discourse, Minerva 56(2), 2017/18, establishes that Bill Gates credited Hilbert's 1900 problems as the inspiration — but Hilbert never used challenge language. The phrase itself derives from sport: the Grand Challenge Cup for rowing at the 1839 Henley Royal Regatta. Explicit grand-challenge definitions entered US science policy through late-1980s computing initiatives, with the National Research Council adopting the language shortly after. Established Its rhetorical work is a substitution. The 1960s and 1970s spoke of “pressing problems” and “crisis problems”; those became “challenges,” which are, in the words of the policy adviser Tom Kalil quoted by Kaldewey, “compelling and intrinsically motivating.” The substitution makes intractable things look tractable. Frontier Kaldewey's term for what follows is the “sportification of science” — competition logic imported into research practice, visible in the DARPA Grand Challenge and RoboCup, producing “self-mobilization and, ultimately, … self-optimization.” His critical conclusion is that most genuine grand challenges — climate, energy security — “remain wicked problems defying clean solutions, yet the discourse assumes feasibility.”
Handwave Now the first cut, and it is the most useful thing on this page because everyone from XPRIZE to Nesta invokes it: the founding case is a myth. Rebekah Higgitt, a professional historian of science and co-author of Finding Longitude, published the archival correction in 2014 under the title There was no such thing as the Longitude Prize. Established There was no single prize. The 1714 Longitude Act “offered a range of sums depending on the accuracy achieved” — a graduated schedule, not a winner-take-all purse. Established The Board of Longitude behaved like a funding agency, not a prize committee. Across 1714–1828, “rewards accounted for only 33% of spending, while overheads (23%), expeditions (15%) and publications (29%) made up the rest.” Two-thirds of the money went into infrastructure — sea trials, instrument work, publication of tables — not into payouts.
Established And Harrison was funded incrementally, like a grantee. He received about £22,000 of £52,534 in total rewards — roughly 42% of all disbursements — through “a number of payments between 1737 and 1764” plus separate Acts of Parliament, with the final £8,750 in 1773 bypassing the Board entirely. Established Nor was the problem solved by one method. “The timekeeping method could never have taken off without the complementary lunar-distance method working in tandem with it”; post-1765 rewards were divided between the two, benefiting Euler and Tobias Mayer among others. Frontier Higgitt's objection to the popular telling is that the Harrison-versus-establishment story “obscures the way that longitude solutions were developed and used.” Handwave So the founding demonstration that “a prize mobilises effort” is, on the record, twenty-seven years of incremental public research funding to a named investigator, alongside a parallel programme on a competing method, misremembered as a lone genius winning a purse.
Established The second cut comes from the only large-sample historical study of inducement prizes, and it says the purse is not the mechanism. Brunt, Lerner and Nicholas, Inducement Prizes and Innovation, Journal of Industrial Economics 60(4), December 2012. The data: Royal Agricultural Society of England prize competitions 1839–1939, 15,032 entrant inventions and 1,986 awards across 98 shows, matched against all British patents granted in the same century — over 900,000 of them. Established A doubling of monetary prizes implies an 11 per cent increase in entrants. An additional gold medal announced implies a 53 per cent increase in entrant counts, and during the 1856–72 prize-rotation period gold medals increased entries by 144%. On quality-adjusted patents — those whose holders paid renewal fees — an additional gold medal during rotation implies a 23 per cent increase. Established Monetary awards covered only about one-third of invention sale prices, and it was “the number of monetary prize awards rather than the monetary amount” that attracted entrants. The authors' own summary: “Medals were more important than monetary awards.”
Frontier Their caveat is theirs, not this brief's, and it matters. It is uncertain whether the observed patenting increase reflects genuine additional innovation or a shift in the propensity to patent, as inventors protected themselves amid increased competition. Frontier The implication for the framing is uncomfortable and testable. If status rather than money is the active ingredient, then grand-challenge framing works as a reputational tournament, and its efficacy depends on the prestige of the convening body rather than on the size of the purse or the importance of the problem. That is a different theory from the one the framing asserts.
Established The third cut is the sponsor's own accounting, and it corrects a claim this site makes elsewhere. DARPA's Report to Congress on the 2005 Grand Challenge, run under the prize authority at 10 U.S.C. § 2374a, records the programme in full. Applications rose from 106 in 2004 to 195 in 2005 — an “84 percent increase,” from 36 states and 3 foreign countries. The 2004 course was 142 miles and the best vehicle managed “approximately 7 miles”; no vehicle finished and the $1 million prize went unclaimed. The 2005 course was 132 miles and four vehicles finished within the ten-hour limit, the Stanford Racing Team winning in 6h 53m 08s for a $2 million prize. Frontier DARPA's own later account says five completed; the contemporaneous report to Congress says four within the limit, and this brief prints both rather than choosing.
Established And the cost line, which is the one everyone omits: DARPA's own expenditure was “approximately $7.8 million, plus the $2 million prize.” Established That is roughly four dollars of agency administration for every dollar of purse, on top of a $1 million 2004 purse that went unclaimed but whose event still had to be staged. Frontier The claim that prizes “are cheap because they pay only for success” is not what the sponsor's own cash flow shows. The prize was cheap relative to a research programme of comparable ambition; it was not cheap relative to its purse, and the formulation is wrong as a description of what the sponsor spent.
Frontier Now the strongest evidence the other way, stated at full strength because this brief owes it. 2004 to 2005: identical terms, twelve months apart, zero finishers to four or five. Nothing else about the field changed that fast. Whatever challenge framing does, it did something in that year, and Nesta's own review reads it exactly that way: “persistence paid off — 2004 failure ($1M) yielded zero completions; 2005 rerun achieved 5 finishes with identical terms.” Established DARPA's own claimed effect — and it is an interested party on itself — is that the competition “attracted thousands of inventors to work in an area important to DoD” and “created a community of innovators… The fresh thinking they brought was the spark that has triggered major advances in autonomous robotic ground vehicle technology.” The 2007 Urban Challenge saw 6 of 11 finish, won by Tartan Racing for $2 million. Frontier No competing explanation for those twelve months has been offered, and this brief does not offer one either.
Established XPRIZE supplies both the leverage claim and its cleanest counter-case. The foundation's own account of the Ansari XPRIZE — an interested party, maximally so, this being its marketing on its founding prize — records a $10 million purse launched in 1996 and won on 4 October 2004 by Mojave Aerospace Ventures' SpaceShipOne, with “$100M invested by competitor teams in R&D”: a claimed 10:1 leverage over the purse. Established The leverage figures are independently reported. Nesta's review records Ansari at “$10 million purse generated $100+ million participant investment” and the Northrop Grumman Lunar Lander Challenge at “$2 million prize spurred $20 million total investment” — both about ten to one. Frontier Whether that induced spending is additional, as opposed to spending that would have happened anyway and merely got labelled, is exactly what nobody has established.
Handwave The same foundation's attribution claim is not evidenced and should be read as advertising. Its page credits the prize with a “$496B+ commercial space industry spurred by competition” and, elsewhere on the same page, with having “ignited the $469B commercial space industry” — two figures $27 billion apart on one page — and with having inspired ventures “like SpaceX, Blue Origin, and others.” Handwave Attributing the global commercial space sector to a $10 million prize is a counterfactual claim with no counterfactual attached, made by the party that benefits from it.
Established And the counter-case is the same foundation's own largest failure. The Google Lunar XPRIZE launched in 2007 with a $30 million purse — $20m grand, $5m second, $5m special — and a deadline extended repeatedly from an original 2012 to 31 March 2018. The field went from “several dozen” teams to five finalists: Moon Express, Hakuto, SpaceIL, Team Indus and Synergy Moon. It ended in January 2018 with no winner and the prize unclaimed, with leadership citing “difficulties of fundraising, technical and regulatory challenges.” Frontier And two finalists said on the record that the prize was not why they were doing it. Moon Express CEO Bob Richards: “The competition was a sweetener in the landscape of our business case, but it's never been the business case itself.” SpaceIL's Ryan Greiss: “SpaceIL is committed to landing the first Israeli spacecraft on the moon, regardless of the terms or status of the Lunar X Prize.” That is the additionality problem in the participants' own words.
Frontier The Millennium Prize Problems are the cleanest available test of framing with nothing else attached, and the result is thin. The Clay Mathematics Institute announced seven problems on 24 May 2000 at the Collège de France at $1 million each. Its stated purposes were to “record some of the most difficult problems with which mathematicians were grappling at the turn of the second millennium” and “to elevate in the consciousness of the general public the fact that in mathematics, the frontier is still open.” Established In twenty-six years, one of seven has been solved — the Poincaré Conjecture, by Grigori Perelman. Six remain open: Birch and Swinnerton-Dyer, Hodge, Navier–Stokes, P versus NP, Riemann, and the Yang–Mills mass gap. Frontier Perelman declined the money; that fact and the solution date were not verified from the Institute's own page in this pass and are carried with a flag rather than as settled. Frontier This is grand-challenge framing with no infrastructure, no programme management, no teams and no downstream customer — just a name, a list and a million dollars — and after a quarter-century it has one solution, produced by a mathematician whose work was already underway and who then refused the prize. Note the Institute's second stated aim, though: public consciousness, not solutions. On its own terms it may well have succeeded.
Established The largest philanthropic instance is Grand Challenges in Global Health, and its scale is not in doubt. Announced by Bill Gates at Davos in January 2003, with an initial $200 million from the Gates Foundation to the Foundation for the NIH, raised to $450 million in May 2005. Fourteen Grand Challenges were announced in October 2003, drawing over 1,000 initial submissions; by August 2004 there were 1,500+ letters of intent from 75 countries and 405+ full proposals through seven peer-review panels, with 43 projects funded by June 2005. Co-funders included the Wellcome Trust at $27.1 million and the Canadian Institutes of Health Research at $4.5 million. Grand Challenges Explorations added $100,000 seed grants twice yearly with up to $1 million follow-on, reaching 3,622 Explorations grants in 117 countries by October 2022. Established And the finding this brief did not expect: the programme's own history page contains no candid admission of failure, no timeline slippage and no lessons learned. It is a scale-up narrative only. That is worth printing precisely because it is the opposite of what an interested party reporting against interest would produce.
Established Which brings the verdict that the evaluation literature reaches about itself, from a source with every reason to say otherwise. Abdullah Gök's Nesta working paper of November 2013 — Nesta runs prize challenges, including the modern Longitude Prize, so its interest runs directly against these findings — reports that “evidence on the effectiveness of prizes is scarce,” that only a handful of rigorous evaluations exist, that most studies are ex ante design assessments rather than ex post outcome measurement, that McKinsey found more than 40% of prizes were not evaluated for their impact at all, and that additionality “remains difficult to verify — counterfactual scenarios unknowable.” Established It adds that “if prizes are poorly designed, managed and awarded, they may be ineffective or even harmful”; that non-monetary incentives matter more than purse size for many participants, converging independently with Brunt, Lerner and Nicholas; that prizes need “achievable and measurable goals” and are unsuitable for basic research; that economic downturns reduce their effectiveness because ex-ante investment becomes prohibitive; and that non-winners still gain — the Progressive Automotive XPRIZE gave non-winning teams “publicity, attention, credibility, access to funds and testing facilities.” Its conclusion: prizes are “complementary under certain conditions,” not substitutes for other instruments. Established A separate conceptual review says the same thing more bluntly: “there has been little empirically-based scientific knowledge on how to design, manage, and evaluate innovation prizes.”
3 · Frontier questions
Established The open question in this subject is a single one wearing several costumes: additionality. Did the challenge cause effort that would not otherwise have occurred? Everything else — entrants, leverage ratios, media attention, community formation — is measurable and largely measured. This is not.
Handwave Is the 10:1 leverage ratio real spending or relabelled spending? Ansari at $10m to $100m+ and the Northrop Grumman Lunar Lander Challenge at $2m to $20m are independently reported and consistent. Frontier What no study establishes is the counterfactual, and the two on-record statements from Google Lunar finalists point the other way: a business case that existed before the prize and a national ambition that would have proceeded regardless. Two data points against a claim is not a refutation; it is the first evidence anyone has offered on the question at all.
Frontier Is status really the active ingredient, and does it generalise? The 53%-per-gold-medal result against 11%-per-doubling-of-purse is the strongest quantitative claim available, from 15,032 entries across a century. Speculative Whether a nineteenth-century agricultural society's medals generalise to a twenty-first-century foundation's press release is entirely untested, and the mechanism would predict something specific and checkable: that a prize convened by a prestigious body outperforms an identical prize from an obscure one. Nobody has run that comparison.
Frontier What happened between 2004 and 2005 at DARPA? This is the sharpest unexplained result in the pack and the strongest evidence for the framing. Identical terms, zero finishers to four or five in twelve months. Speculative The competing explanations — learning from a public failure, team consolidation, sensor cost curves, the sheer growth in applications from 106 to 195 — have not been separated by anyone. A decomposition of that twelve months would be the most valuable single piece of work anyone could do on this subject.
Frontier Does pure framing with no programme attached do anything? The Millennium Problems are the closest thing to a controlled test: a name, a list, seven million dollars, no infrastructure. One solved in twenty-six years, and the solver refused the money. Speculative The counter-reading is that the Institute's own second aim was public consciousness rather than solutions, in which case the instrument should be evaluated on attention and not on the solution rate — and on that measure nobody has measured it either.
Frontier And a question about the frame itself rather than the instrument. Kaldewey's reading is that grand-challenge language is identity work: its function is to make wicked problems feel tractable, by substituting “challenge” for “crisis.” Speculative If that is right, then the framing's real output is institutional legitimacy and researcher self-mobilisation rather than solved problems, which is a coherent theory nobody has attempted to test against outcomes. Handwave It is also unfalsifiable as usually stated, and saying so is part of taking it seriously.
4 · Technological bottlenecks
Established The bottleneck is the counterfactual, and it is structural rather than practical. To know whether a challenge caused effort, you need to know what the same actors would have done without it. Established Gök's phrasing is exact and comes from an organisation that runs prizes: additionality “remains difficult to verify — counterfactual scenarios unknowable.” That is not a gap in the literature. It is a property of the object.
Established Second bottleneck: nobody is required to evaluate. More than 40% of prizes were not evaluated for their impact at all, and most of what exists is ex ante design assessment rather than ex post outcome measurement. Frontier A sponsor who runs a prize, declares success in a press release and moves on has satisfied every obligation anyone imposes on them, and the resulting literature is dominated by material produced by parties with an interest in the answer.
Frontier Third: the outcome measure is contested and often unstated. Entrants, patents, participant spending, media reach, community formation, solved problem, public consciousness — sponsors switch between these depending on what the result supports. Established The Clay Institute is unusually honest here in stating both of its aims up front, which is why it is the easiest case to evaluate fairly and the one whose apparent failure is least damning.
Frontier Fourth, a measurement problem inside the best evidence available. Brunt, Lerner and Nicholas measure patenting, and say themselves they cannot separate genuine additional invention from a shift in the propensity to patent under competitive pressure. Speculative The single best empirical result in this field has an identification caveat its own authors state, and every downstream use of it that omits the caveat is overstating it.
Established And fifth: the founding case cannot bear the weight put on it, which removes the historical base the whole instrument was sold on. Longitude was graduated rewards inside what was functionally a research agency spending two-thirds of its money on overheads, expeditions and publications. Handwave An instrument whose canonical precedent turns out to be a different instrument has no demonstrated track record before the twentieth century at all.
5 · Research dependencies
Established This brief depends on the DARPA material owned elsewhere on the site, and its one intrusion into that territory is deliberate and should be read as a correction. The ARPA model, the programme-manager mechanism, the Heilmeier catechism and the preconditions for replication all belong to the DARPA study. Established What this brief owns is the Grand Challenge as a prize instrument, and the specific claim it corrects is that prizes “are cheap because they pay only for success.” DARPA's own report to Congress records approximately $7.8 million of agency expenditure plus a $2 million prize. Frontier That sentence should be qualified by cross-reference rather than left standing.
Established It depends on Public Policy Foresight for the shape of its own finding, not for its content. That brief documents a field that declined to define an outcome measure for itself; this one documents an instrument whose sponsors mostly declined to evaluate it. Two independent policy domains arriving at the same structural absence is more informative than either alone, and the cross-link is worth following in both directions.
Established It depends on one empirical paper for its central quantitative claim. The 11%-versus-53% result is the only large-sample historical estimate of what actually attracts entrants, and there is no replication of it. Frontier Single-source dependency on the load-bearing number, with an identification caveat the authors state themselves.
Established It depends on an interested party for its most damaging general finding, and that is a feature. Nesta runs prizes, including the modern Longitude Prize, and published that the evidence base for prizes is scarce and that additionality is unverifiable. Frontier Interest running against a finding raises its weight, and this is the clearest instance of that rule anywhere in this category.
Frontier And it depends on second-hand reporting for two named criticisms it carries. Anne-Emanuelle Birn in The Lancet (2005) and Laurie Garrett in Foreign Affairs (2007) reach this brief through an encyclopaedia entry rather than from the originals. Neither should be treated as read, and both are listed so the route is visible.
6 · Required experiments
Established The experiments this subject needs are unusually feasible, because prizes are already run in numbers and the design variation is free. Almost nothing in the list below requires a new institution — it requires sponsors to randomise something they are already choosing arbitrarily.
Frontier One: randomise the non-monetary component. The strongest claim in this field is that medals beat money. A sponsor running repeated rounds could vary recognition tiers while holding the purse constant, which is exactly the variation Brunt, Lerner and Nicholas exploited historically by accident. Speculative This is a within-programme design that any large challenge sponsor could run at negligible cost, and none has.
Frontier Two: randomise the convener. If status is the mechanism, an identical challenge announced by a prestigious body and by an obscure one should draw different fields. Speculative That is the sharpest available test of the reputational-tournament theory and it has never been attempted, presumably because no prestigious body wants to discover that its name is the product.
Established Three: collect entrant counterfactuals at registration. Ask every entrant, at the point of entry, what they would be doing otherwise and with what funding. Frontier It is self-reported and weak, and it is infinitely more than the zero evidence that currently exists on additionality. The two Google Lunar statements that undercut the induced-effort model came out incidentally, in press interviews, after the prize had failed.
Frontier Four: decompose DARPA 2004 to 2005. Team-level data on continuity, spending, personnel and component costs across the two years would separate the framing effect from learning, consolidation and falling sensor prices. Speculative The data existed at the time and the decomposition has never been published. It is the highest-value single study available in this subject.
Established Five, the retrieval this brief could not make and names as a live task: a published self-evaluation of Grand Challenges in Global Health. The programme's own history page carries no candid assessment; a decade retrospective in a major journal is paywalled and its contents are not quoted here; and the widely repeated Gates admission of naivety about development timelines could not be sourced to any fetched document and is therefore not quoted anywhere in this brief. Frontier This was the expected model case for an interested party reporting against interest, and it is currently a hole.
7 · Engineering requirements
Established The engineering here is the design of the instrument, and the evidence supports a short list of specifications rather than a theory.
Established A prize needs an achievable and measurable goal, and is unsuitable for basic research. That is Gök's finding and it is not seriously contested. Frontier The Millennium Problems are the natural experiment on the boundary: seven of the hardest open questions in mathematics, a million dollars each, and one solved in twenty-six years by someone who was working on it anyway.
Established Recognition should be designed as carefully as the purse, and probably more carefully. Doubling the money bought 11% more entrants; one more gold medal bought 53%, and 144% during the rotation period. Frontier The design implication is specific: more award tiers rather than a bigger top prize — it was the number of monetary awards, not the amount, that attracted entrants, with monetary awards covering only about a third of invention sale prices.
Established Budget the administration, not the purse. DARPA spent about $7.8 million to award $2 million, and staged a 2004 event whose $1 million purse went unclaimed. Frontier A sponsor budgeting on the purse alone is budgeting about a fifth of the cost, and the “pay only for success” framing is precisely what causes that error.
Established Design for the non-winners, because the measurable benefits accrue to them. The Progressive Automotive XPRIZE gave non-winning teams “publicity, attention, credibility, access to funds and testing facilities.” Frontier If the field's most defensible benefit is community and capability formation among entrants who lose, then a design that maximises entry and visibility outperforms one that maximises the top prize — which is the same conclusion the medal result reaches from a different direction.
Frontier Expect to repeat, and budget for a failed round. The single strongest pro-framing datum in the record is a rerun: 2004 to 2005, identical terms, zero to four or five. Speculative A sponsor who cancels after an unclaimed first round may be cancelling one year before the result — and a sponsor who reruns indefinitely may be doing what the Google Lunar XPRIZE did, extending a 2012 deadline to 2018 and ending with no winner. Nothing in the evidence distinguishes the two situations in advance.
8 · Adjacent technologies
Established The seam with the DARPA material is the one this brief deliberately crosses, once, and it should be read as a correction rather than an encroachment. The ARPA model, the programme-manager mechanism, the Heilmeier catechism and the conditions for replicating the agency all belong to the DARPA study, which this brief does not restate. Its single intrusion is the $7.8 million administrative figure from DARPA's own report to Congress, which qualifies the claim that prize competitions “are cheap because they pay only for success.” Frontier That claim is right that prizes mobilise effort grant funding does not; it is wrong about the sponsor's cash flow, and the correction belongs on the record rather than in a footnote.
Established The second seam is with Public Policy Foresight, and it is a shared shape rather than a shared subject. That brief owns the evaluation of government futures machinery and the documented absence of impact evidence in it. This brief owns the challenge framing as an instrument, and arrives at the same structural finding by an independent route: an instrument in wide use whose sponsors mostly did not measure whether it worked. Frontier Two domains, no common authors, the same absence — which is weak evidence that the absence is a property of policy instruments generally rather than of either field.
Established Third, the one-goal one-deadline model is a different instrument and should not be blended in. Apollo and the Manhattan Project are directed programmes with a customer, a budget and a chain of command; a prize is an open call with a specification and no employment relationship. Frontier Conflating them is the most common error in grand-challenge advocacy, because it lets the record of the directed programmes stand in for the record of the prizes, which is much thinner.
Established Fourth, the funding-institution seam. Scientific Governance Models and Scientific Advisory Institutions own how research money is allocated and disciplined; this brief supplies them the finding that the best-documented historical alternative to grant funding turns out, on inspection, to have been grant funding with a graduated reward schedule attached. Frontier The Board of Longitude spent 67% of its money on overheads, expeditions and publications, which is a research agency's budget.
Speculative And a seam with the measurement problem that runs through this whole category. The historical briefs here share a structure: a widely held claim about how things work, and an evidence base that either tests it or turns out never to have been assembled. Handwave Treating that recurrence as a method rather than a coincidence is an editorial judgement, and it is stated as one.
9 · Institutional requirements
Established The first institutional finding is that the founding institution of this instrument was not running the instrument. The Board of Longitude spent 33% of its money on rewards and 67% on overheads, expeditions and publications across 1714–1828; it paid Harrison in instalments between 1737 and 1764; and his final £8,750 in 1773 came by Act of Parliament, outside the Board entirely. Handwave Every modern sponsor invoking Longitude as proof of concept is invoking a research funding agency, and the fact that nobody in the chain checked is the same failure this category keeps finding.
Established Second, the instrument requires statutory authority in government hands and that authority shapes it. DARPA ran its Grand Challenge under 10 U.S.C. § 2374a and reported to Congress on it, which is why the $7.8 million administrative figure exists at all. Established A prize sponsor that reports to a legislature discloses its costs; a foundation that does not, does not. Frontier The comparison across the sources in this brief is stark: the agency accountable to Congress published its administrative overhead against its own interest, and the foundation published a $496 billion attribution and a $469 billion attribution on the same page.
Established Third, no institution is obliged to evaluate a prize, and the consequence is measured. More than 40% of prizes were never evaluated for impact. Frontier The evaluation gap is not an oversight by researchers; it is the predictable output of a system in which the only party with the data is the party with an interest in the answer.
Established Fourth, the one institution that reported honestly against interest is the one that runs prizes. Nesta published that the evidence is scarce, that most studies are ex ante, that additionality is unverifiable and that poorly designed prizes “may be ineffective or even harmful.” Established That is an organisation documenting the weakness of its own instrument, and it raises the weight of the finding rather than lowering it. Frontier It is also, on the evidence assembled here, unusual.
Frontier And fifth, the institutional design question the evidence actually poses. If status is the active ingredient, then the scarce resource being allocated is the convener's prestige, not its money. Speculative That would make grand challenges a mechanism by which established institutions convert reputation into other people's labour, which is neither obviously good nor obviously bad and is a different object from the one the policy literature describes. Handwave Nobody has tested it, and stating it as anything stronger than a hypothesis would be doing exactly what this brief criticises.
10 · Ethical & societal considerations
Established The first ethical issue is the transfer of risk from sponsor to entrant, and the numbers make it concrete. A prize pays only the winner. The Ansari purse was $10 million against a reported $100 million+ of participant investment; the Northrop Grumman challenge, $2 million against $20 million. Frontier Ten to one leverage is also nine parts unrecovered spending by people who did not win, and the Google Lunar XPRIZE ended with $30 million unclaimed after eleven years and a field that had shrunk from several dozen teams to five.
Frontier Second, the honest counterweight, because the evidence supports it. Non-winners demonstrably gain: the Progressive Automotive XPRIZE gave losing teams “publicity, attention, credibility, access to funds and testing facilities.” Speculative Whether those gains are worth the unrecovered spending is a question nobody has attempted to price, and the answer almost certainly varies by whether an entrant was a venture-backed company or a graduate student.
Handwave Third, the attribution claims are an ethical problem and not merely an accuracy one. A foundation crediting a $10 million prize with a $496 billion industry — and a $469 billion one, on the same page — is making an unfalsifiable claim that recruits future entrants and future sponsors. Frontier The people who bear the cost of an overstated track record are the next cohort of teams that spend nine dollars for every one on offer.
Established Fourth, the substantive criticism of the framing itself, from named critics on the record. Anne-Emanuelle Birn argued in The Lancet in 2005 that the initiative was “weak” for focusing “too narrowly on the power of science” and neglecting economic and social factors, malnutrition being “political and economic” rather than technical. Laurie Garrett argued in Foreign Affairs in 2007 that targeting specific high-profile diseases “is not enough to improve public health, as education and an all-disease health system are needed.” Frontier Both reach this brief at second hand and neither should be treated as read. Speculative The structural version of their point is Kaldewey's: a frame whose function is to make problems look tractable will select for the problems that can be made to look tractable.
Speculative And fifth, the ethics of the myth. Longitude is invoked constantly and is not what it is invoked as. Handwave A sponsor who cites it in good faith is passing on an error; a sponsor who cites it after being shown Higgitt's archival figures is doing something else — and the distinction is worth naming, because the correction has been in print since 2014 and the invocation has not stopped.
11 · Civilizational implications
Established The durable finding is the split, and it is worth stating in one sentence. Challenge framing reliably attracts entrants, and whether it produces innovation that would not otherwise have happened is essentially unmeasured. Established Entry effects are documented at 11% per doubling of purse, 53% per additional gold medal, and 106 to 195 applications in a year at DARPA. Additionality has one large-sample study with a stated identification caveat, two on-record participant statements against it, and an evaluation literature that describes its own evidence as scarce.
Frontier Second, if status is the active ingredient, the implications are large and mostly unwelcome. It would mean the effectiveness of a grand challenge depends on who is convening rather than on what is being solved, and that a society's capacity to mobilise effort by naming a problem is a function of how much prestige its institutions still hold. Speculative That would make the instrument's efficacy a lagging indicator of institutional trust — an interesting and completely untested claim.
Established Third, the sportification observation deserves to be carried forward rather than left as a critique. Kaldewey's point is that competition logic imported into research produces “self-mobilization and, ultimately, … self-optimization,” and that the substitution of “challenge” for “crisis” makes intractable things look tractable. Frontier Most real grand challenges — climate, energy security — remain wicked problems defying clean solutions, and the discourse assumes feasibility. Speculative A civilization that can only mobilise around problems that can be made to look winnable has a systematic blind spot shaped exactly like its hardest problems.
Frontier Fourth, the DARPA result is the strongest reason not to dismiss the instrument, and it should be held at full strength. Zero finishers to four or five, twelve months, identical terms. Established Whatever happened there is real and nobody has explained it away. Speculative If the effect is repeatable, the implication is that a public failure followed by an immediate rerun is a more powerful instrument than either the failure or the prize alone — which is not how sponsors behave.
Speculative And fifth, the meta-observation this brief keeps arriving at. An instrument in wide use for decades, with hundreds of instances, and more than 40% never evaluated. Handwave The civilizational finding is not about prizes at all: it is that we deploy policy instruments at scale and decline to build the measurement that would tell us whether to keep deploying them — and the same absence turns up independently in the foresight literature, which is two domains rather than a pattern.
12 · Timelines
These horizons track whether the additionality question gets answered, not whether prizes keep being run — they will be run regardless:
- 10 yr: Frontier Expect prize volume to keep rising and evaluation to keep lagging, because no sponsor faces a requirement to evaluate and the party with the data is the party with an interest in the answer. Speculative The cheap studies either happen or do not: a decomposition of DARPA 2004 to 2005, entrant counterfactuals collected at registration, and a randomised recognition tier inside a repeated programme. Frontier Expect the Longitude myth to survive the correction, since it has already survived a decade of it in print.
- 25 yr: Speculative The plausible split is that entry effects become well measured, because entrants are countable, and additionality stays unmeasured, because the counterfactual is unknowable by construction rather than by neglect. Speculative If any large sponsor randomises the non-monetary component, the reputational-tournament theory becomes testable and the field acquires its first real design evidence. Handwave Whether a sponsor does that depends on whether one is willing to learn that its name is the product.
- 50 yr: Speculative At this range the Millennium Problems become the most informative object in the subject: seven problems, one solved in twenty-six years, and a solution rate that will be readable as a long-run measurement of what pure framing with no programme attached achieves. Speculative If four or five remain open at seventy-five years, that is about as clean a null as this instrument will ever generate. Handwave It is also seven observations, which is not a sample.
- 100 / 250+ yr: Handwave Beyond useful forecasting. Handwave The only durable observation is the one this brief opens with: the instrument's founding case was misremembered within a couple of generations, and the misremembering is what got copied forward. Expect that to be the dominant process at this horizon, on this subject as on every other.
13 · Technology tree & dependencies
- Depends on Nothing on this map produces a result this brief is waiting for. The evidence that exists — a century of agricultural prize competitions, a sponsor's report to Congress, an archival correction to the Longitude story, a foundation's own failed lunar prize — is already published, and what is missing is a study design nobody has run rather than a finding another brief will supply. No typed depends-on edge is claimed.
- Requires (not on this map) Two things sponsors could supply tomorrow and do not, and between them they are the entire reason this instrument is unevaluated after two centuries. First, counterfactual data collected from entrants at the point of entry: what each team would be doing without the challenge, on what funding, on what timetable. Additionality is the whole question — Nesta's own review calls counterfactual scenarios unknowable — and the only evidence that currently bears on it arrived by accident, in press interviews after the Google Lunar XPRIZE failed, when Moon Express's chief executive said the competition “was a sweetener in the landscape of our business case, but it's never been the business case itself” and SpaceIL said it was committed to its mission “regardless of the terms or status” of the prize. Self-reported intentions at registration are weak evidence; they are also infinitely stronger than nothing, which is what exists. Second, an evaluation requirement attached to publicly and philanthropically funded prizes, of the kind that already attaches to most other instruments spending the same money. More than 40% of prizes have never been evaluated for impact, most published studies assess design before the fact rather than outcomes after it, and the only party holding the data is the party with an interest in the answer. Neither requirement is a research result. Both are conditions a funder could impose in a grant agreement, and until one does, the honest position on whether challenge framing works will remain that entry rises and nobody knows what else does.
- Enables This brief supplies a priced, sourced correction to a claim the site makes elsewhere — that prize competitions are cheap because they pay only for success — and an evidence inventory that any brief invoking challenge framing can be checked against. No typed enabling edge is claimed, because a correction and an evidence inventory are inputs to how other arguments are read rather than prerequisites for anything being built.
- Adjacent The economics of innovation, which supplies the inducement-prize literature and the additionality problem; the history and sociology of science policy, which supplies the genealogy and the sportification thesis; programme evaluation methodology, which supplies the counterfactual machinery this subject lacks; and within this map Public Policy Foresight, Scientific Governance Models, Scientific Advisory Institutions, Long-Term Institutions and Civilizational Planning.
14 · Common misconceptions & speculative claims
Handwave “The Longitude Prize proves that prizes work.” There was no Longitude Prize in the sense intended. The 1714 Act “offered a range of sums depending on the accuracy achieved.” Across 1714–1828 rewards were 33% of Board spending, with overheads at 23%, expeditions at 15% and publications at 29%. Harrison was paid in instalments between 1737 and 1764, receiving about £22,000 of £52,534, and his final £8,750 in 1773 came by Act of Parliament outside the Board. The problem was not solved by one method: post-1765 rewards were split with the lunar-distance approach. Established The founding demonstration of the instrument is twenty-seven years of incremental public research funding to a named investigator, misremembered as a lone genius winning a purse — and the correction has been in print since 2014.
Handwave “Prizes are cheap because they pay only for success.” Not what the sponsor's own accounting shows. DARPA's report to Congress records approximately $7.8 million of agency expenditure plus the $2 million prize for the 2005 Grand Challenge — about four dollars of administration for every dollar of purse — on top of a $1 million 2004 purse that went unclaimed but whose event still had to be staged. Established This is an interested party disclosing a cost against its own interest, which is the strongest kind of evidence available here. Frontier The defensible version of the claim survives: a prize is cheap relative to a research programme of comparable ambition. It is not cheap relative to its purse, and budgeting on the purse alone budgets about a fifth of the cost.
Frontier “Bigger purses attract more effort.” True and much weaker than assumed. In the only large-sample historical study — 15,032 entrant inventions across 98 shows, matched to over 900,000 patents — doubling the money bought 11% more entrants and one additional gold medal bought 53%, rising to 144% during the rotation period. It was “the number of monetary prize awards rather than the monetary amount” that mattered, and monetary awards covered only about a third of invention sale prices. Established The authors' own conclusion is that “medals were more important than monetary awards.” Frontier Their caveat belongs with the result: they cannot separate genuine additional invention from a shift in the propensity to patent under competitive pressure.
Handwave “The Ansari XPRIZE created the commercial space industry.” The foundation's own page credits the prize with a “$496B+” industry and, elsewhere on the same page, with igniting a “$469B” one — two figures $27 billion apart — and with inspiring SpaceX and Blue Origin. Handwave No counterfactual is offered and none exists. Established What is independently supported is the leverage ratio: about 10:1 in participant spending, corroborated for Ansari and for the Northrop Grumman Lunar Lander Challenge by an outside reviewer. Frontier Whether that spending is additional is precisely what nobody has established.
Frontier “A big enough prize will get the thing built.” The Google Lunar XPRIZE was $30 million, ran from 2007, extended its deadline repeatedly from 2012 to 31 March 2018, went from several dozen teams to five finalists, and ended in January 2018 with no winner. Established And two of its finalists said publicly that the prize was not their reason. Moon Express: “a sweetener in the landscape of our business case, but it's never been the business case itself.” SpaceIL: committed “regardless of the terms or status of the Lunar X Prize.” That is the additionality problem stated by the participants themselves.
Frontier “Naming a hard problem and attaching money to it mobilises the field.” The Millennium Prize Problems are the cleanest test: seven problems, $1 million each, announced 24 May 2000, no infrastructure, no programme management, no teams, no customer. One solved in twenty-six years, and the solver declined the money. Frontier But read the Institute's own stated aims before scoring it: the second was to “elevate in the consciousness of the general public the fact that in mathematics, the frontier is still open,” and on that measure it may well have succeeded. On the framing under test it did not mobilise effort in any way that shows up in the solution rate.
Handwave “Prizes work for anything.” The evaluation literature says the opposite and is specific: prizes need “achievable and measurable goals” and are unsuitable for basic research; they are “complementary under certain conditions,” not substitutes for other instruments; and “if prizes are poorly designed, managed and awarded, they may be ineffective or even harmful.” Established That is an organisation which runs prize challenges publishing the limits of its own instrument.
Frontier And the mirror misconception: “prizes are useless.” Not supported and this brief refuses it. DARPA 2004 to 2005 is real — identical terms, zero finishers to four or five in twelve months, applications up 84%, and no competing explanation offered. The 10:1 leverage ratios are independently reported. Non-winners demonstrably gain publicity, credibility, funding access and testing facilities. Established The honest terminal position is the split: entry effects are established, additionality is unmeasured, and the two should never be reported as one finding.
Established Four things this brief deliberately does not print, so their absence is not read as oversight. Bill Gates's widely quoted admission of naivety about development timelines — it could not be sourced to any fetched document and is not quoted here. The contents of the decade retrospective on Grand Challenges in Global Health — metadata confirmed, body paywalled. Birn's and Garrett's criticisms as read — both reach this brief at second hand. Frontier And the empirical study underlying one cited review, for which a bibliographic query returned stale, wholly unrelated content in this pass; no citation for it has been constructed, and constructing one would be exactly the failure the Longitude and Interavia cases in this category are about.