1 · Concept overview

This brief is about the design of the funder, not the design of the judgement inside it. Who gets money, how much, for how long, on what contract, chosen by whom, terminable on what grounds, and at what fixed cost to the system as a whole. The layer beneath — peer review as a scoring and ranking instrument, its reliability, blinding, registered reports, replication policy, integrity enforcement, and lotteries as a within-funder allocation rule — is owned by Scientific Governance Models and is used here rather than re-derived.

The framing under test is a sentence with two halves joined by and: that how science is funded determines what science gets done, and that the alternatives are known. They come apart, and the split is the finding. The first half has one well-identified demonstration and several weaker corroborations. The second half does not survive contact with the record at all. Every alternative on the reformer's list — block grants, agencies with empowered programme managers, focused research organisations, prizes, advance market commitments, people-not-projects awards, partial lotteries — is named, described, advocated, and in almost every case running somewhere with real money. For none of them is there a published comparison of research outcomes against the mechanism it is meant to replace.

2 · Current scientific position

Frontier The one place where “funding determines what science gets done” is demonstrated rather than asserted is the comparison of Howard Hughes Medical Institute investigators against similar NIH-funded scientists. Azoulay, Graff Zivin and Manso compared investigators appointed in 1993–1995 against NIH-funded early-career prize winners using propensity-score weighting and semiparametric difference-in-differences. Seventy-three investigators were appointed; after imposing common support the estimation sample is 69 treated and 348 controls, 417 scientists, with 4 treated and 45 controls dropped. Frontier The estimated effects rise monotonically through the citation distribution: +39% publications overall, +31% in the top 25% of citations, +55% in the top 5%, +97% in the top 1%. Established And in the same design, the same scientists produced 35% more articles falling below the vintage-adjusted citation floor of their own least-cited pre-appointment work, with roughly 10% lower keyword overlap between pre- and post-appointment work and output cited by a more diverse set of journals. The design bought more hits and more misses at once. That is the load-bearing result in this brief, and note what it is not: it is a different output distribution, not a better one.

Frontier Two numbers from the same paper are routinely dropped and change how it should be read. Treated investigators filed 3.2 NIH applications on average between 2003 and 2008 against 5.1 for controls, and the applications they did file received worse priority scores from NIH panels. The treatment reduced engagement with the mechanism it is being compared against, and the comparison institution scored the same people down. Speculative Whether that reflects panel conservatism about exploratory work or genuine lower quality in the marginal application is not separable in this design. Frontier Career outcomes run the same direction and are weaker evidence: 33% of treated investigators were elected to the National Academy of Sciences against 4.1% of controls, and they trained 1.13 early-career prize-winning trainees each against 0.23 — end-of-career levels on a group selected partly on promise, not causal estimates of the same quality.

Established What the contract actually is, institutionally, is the part this brief owns. An initial seven-year term, renewable indefinitely on successful scientific review, worth roughly $11 million per investigator across the term covering salary, benefits, research budget and equipment. The institute reports more than 250 investigators and committed more than $300 million to a cohort of 26 new appointments, under the slogan “people, not projects”. Frontier The operative features are the horizon and the failure tolerance, not the money: a seven-year renewable term with review on scientific merit rather than on delivery against a proposal is a different contract from a three-to-five-year project grant, and it is the contract, not the institute, that the result is about. Mark the source of the programme description — it is the funder describing itself — while the comparison is by outside economists.

Frontier On block against competitive funding, there is one serious study and its primary hypothesis failed. Wang, Lee and Walsh surveyed Japanese research-active faculty on publications from 2001–2006: 2,081 responses at a 27% response rate, 1,399 professors in the analysis, 9,558 sampled publications across 22 fields, with novelty measured as the atypicality of referenced journal pairs. They predicted competitive funding would suppress novelty. It did the opposite: competitively funded projects showed +0.35 to +0.60 higher novelty. Frontier The status-contingency hypothesis was supported and is the interesting half: the positive effect reverses for low-status researchers, with interaction coefficients of −0.94 for junior researchers and −0.95 for female researchers — the latter on 5% of the sample — and −0.64 but not significant for researchers at peripheral universities. Established The composition figures matter for reading any of it: internal block funds supplied a mean 40.92% of project funding and competitive grants 32.54%. This is a comparison of funding mixes inside one national system, not of pure block against pure competitive regimes. Handwave Every public argument that block funding produces more or less adventurous science than project funding is currently doing its work by assertion, on this base.

Established The ARPA model is unusually well described and unusually poorly evaluated, and the gap between those two facts is the point. The canonical statement gives four features argued to be complements — organisational flexibility, bottom-up programme design, discretionary project selection and active project management — so that partial adoption should be expected to underperform, with DARPA characterised as roughly $3 billion of extramural research funding overseen by about 100 programme managers serving single terms of three to five years. Established Goldstein and Kearney supply the only systematic measurement of the discretion, on 234 completed ARPA-E projects with end dates on or before 31 December 2015, out of 471 initiated, alongside 10,227 concept papers and 2,335 full applications: 85% of projects had timeline changes, 43% budget modifications, 45% milestone additions or deletions, and 70% experienced at least one programme-director transition. On selection, roughly 49% of selected applications were promoted — chosen despite falling below a hypothetical peer-review score cutoff — with mean review score predicting selection at about a 20% increase in probability per point against 47% under a cutoff model.

Frontier On outcomes the same study reports 43% of projects publishing, 44% filing patent applications and 13–16% showing market engagement, with milestone-modified projects 18% more likely to show market engagement and projects with a programme-director change 15% more likely to file patents — the latter contradicting the authors' own hypothesis that turnover would hurt. Established And the authors refuse the causal reading in terms worth adopting verbatim: because milestone changes are not applied randomly between projects, they cannot make a claim about the impact of active management on project performance, and every comparison is internal to an agency that practises active management everywhere, so the contrast is degree rather than presence. Established The agency's own cumulative reporting through September 2023: 1,465 projects funded since 2009, $3,508.2 million appropriated, 150 new companies formed, 27 companies acquired, merged or taken public with $21.9 billion in total market valuations, $11.8 billion in private follow-on funding across 218 project teams, 1,073 US patents, 7,047 journal articles and 323 reported patent licences — described by the agency as key early indicators, with no discussion of what they miss and no comparison group. Established The honest limitation is stated by the model's own advocates and is about windows rather than about the agency: it is impossible to measure the incidence of one-in-a-thousand ideas on a timescale relevant to political decisions about programme authorisation. Projects run about three years; the transformations that justify the model take decades.

Established Focused research organisations are the newest model and have no outcome evidence at all — a statement, not a criticism. The design is a time-bound non-profit team of 10–30 full-time staff, $20–50 million, three to seven years, aimed at problems described as too complex or infrastructure-heavy for a single academic lab or a company, producing public goods. The incubator reports almost a dozen launches over four years and estimates that 100 to 200 problems of that shape currently exist. Speculative The evidence that the model works is two named technical results in an essay whose four authors all work for the incubator: an imaging system, and a reported thirty-fold increase in proteomics sample throughput. No adoption metrics, no comparison group, no completed-organisation outcome accounting, and the public page on end-states lists possibilities without naming a completed case. Frontier The against-interest content in the same source is worth more than the advocacy: rigid milestone-setting failed and was replaced with adaptive intermediate goals and dynamic quarterly objectives, and product-development demands were underestimated. That is a direct concession that the milestone-contract feature borrowed from the ARPA model did not transfer.

Frontier The adjacent experiment is the standing institute funded to skip the grant system entirely. Arc Institute launched in 2021 with $650 million, funding core investigators for up to 20-person labs over eight years with no requirement to apply for external grants, plus $1 million over five years innovation awards and $100,000 one-year ignite awards, in partnership with three universities. Established Its reported outputs to date are scientific rather than evaluative. Handwave Any claim that these organisations outperform the grant system is currently unfalsifiable in practice — they are younger than the outcome lag of the work they do, and nobody has specified the comparison that would settle it. Saying the evidence is thin overstates it; there is no outcome evidence.

Established Prizes have the longest measured record and it is thinner and stranger than the advocacy suggests. The best-identified historical evidence covers the Royal Agricultural Society of England, 1839–1939: 15,032 entrant inventions, 1,986 awards across 98 shows, matched against over 900,000 British patents with 20,542 examined for renewal fees. On entry, doubling monetary prizes increased entrants by 11% and each additional medal increased expected entrants by about 12%, with gold medals raising entries by 144% during a rotation period. On patenting, 22% of prize winners patented their exhibited inventions against 17% of non-winning entrants, an additional medal yielding an 8% increase, a gold medal 12–16%, and 23% for renewed patents. Established The finding that most cuts against how prizes are sold is that medals outperformed money. Prize cash covered only about a third of an invention's sale price; certification and free advertising did the work. Displacement tests were negative — only 20.9% of repeat entrants switched technology category — so the prizes appear to have added effort rather than redirected it. Frontier The authors cannot exclude a shift in propensity to patent among publicity-seeking inventors, and unpatented innovation stimulated by prizes is unmeasured.

Established The modern prize record is weak, and its weakness is documented by a prize operator. A review by an organisation that runs prizes states plainly that the evidence on their impact “is scarce”, records a finding that over 40% of prizes received no impact evaluation at all, and notes that evaluated prizes used inconsistent ad hoc methods. It also collects the failures: a substantial number of unclaimed prizes; the 2004 DARPA Grand Challenge, which no participant completed, succeeding only after redesign in 2005; and a super-efficient refrigerator programme whose winning product did not attract consumers. It reports an estimate that a $10 million spaceflight prize induced about $100 million of investment while noting the calculation's transparency problems, and a finding that career and social motivations outweighed cash for online problem-solvers. Frontier Interest running against interest: a prize operator publishing that the evidence for prizes is scarce and that a large share are never evaluated. Frontier The largest recent prize completed its run — an £8 million antimicrobial-resistance prize launched in 2014 and awarded in 2024 for a system identifying a urinary bacterial infection in about 15 minutes and the correct antibiotic in about 45. The announcement is by the prize's own operator, lists achievements and states no limitations; the bare fact worth carrying is the timescale: ten years from launch to award.

Frontier The pull mechanism with the largest measured claim is the advance market commitment, and the deflation comes from the economists who designed it. The $1.5 billion pilot for pneumococcal conjugate vaccine is associated with more than 150 million children immunised, annual distribution above 160 million doses by 2016, and an estimated 700,000 lives saved. Established The same authors record that the target was “a technologically close target” — the vaccines already existed and were in late-stage trials, so the mechanism incentivised manufacturing capacity rather than breakthrough research — that India, the most populous eligible country, did not participate until 2017 and then in five states, and that “we lack a valid counterfactual.” An advance market commitment that worked is on the record. One that induced an invention is not.

Established Concentration is the best-measured feature of the existing system and it is measured by the agency's own extramural leadership. Across FY1985–2020, the top 1% of research-project-grant principal investigators received 8% of funds in 1998, rising to nearly 10% by 2020 — a two-point shift worth about $420 million, roughly 800 grants. Median FY2020 funding was $4.8 million for the top centile against $0.4 million for everyone else, and 80.2% of the top centile held two or more awards against 32.7% of the rest. Established The comparison the authors themselves draw is the one to carry: from 1995 to 2019 the top centile of the US population went from 14.3% to 18.7% of income, a 31% relative rise, while the top centile of principal investigators went from 8.3% to 10.8% of funding, a 30% relative rise. Research funding is less concentrated than national income and has been concentrating at almost exactly the same rate. Established A Theil decomposition in the same paper finds that for every grouping tested — career stage, gender, race, degree, organisation, region, state — within-group inequality exceeds between-group inequality, so any reform framed entirely around between-group gaps is aimed at the smaller component. Established At organisation level the top 10% of organisations received about 70% of funds. The interest runs against the publisher again: agency leadership publishing that their own distribution has concentrated, and calling its racial and gender components unacceptable inequalities.

Frontier Institutional concentration is extreme and roughly stable, which is not what the reform literature assumes. Federal science and engineering support to higher education rose from $31.2 billion to $49.0 billion across 1,110 institutions between FY2014 and FY2023; the top 10 institutions took $7.2 billion (22.7%) in FY2014 and $9.9 billion (20.1%) in FY2023; the top 100 took $38.4 billion, 78.4% of the total, in FY2023. Established Concentration at the very top fell slightly in share while rising sharply in dollars, with pandemic supplemental appropriations explaining much of the movement and rankings returning toward the pre-pandemic pattern from FY2022. The claim that institutional concentration is rapidly worsening is not supported by this series. Frontier The efficiency argument against concentration is a proposal rather than a finding: two agency analyses locating a productivity sweet spot near $400,000 in total annual costs per investigator, with output tapering above $300,000 of annual direct costs, and incremental returns at $800,000 and at $200,000 estimated at about 25–40% of returns at the sweet spot, supporting a proposed floor and cap that would free about $4.22 billion and fund roughly 10,500 additional investigators, a 38% increase in the funded pool. Speculative That calculation assumes the marginal investigator added at the sweet spot would produce at the sweet-spot rate, which is exactly the extrapolation a diminishing-returns curve drawn across investigators the current system selected cannot license.

Established The system's fixed cost is well measured and the one attempt to cut it made it worse. A 2018 workload survey of 11,167 principal investigators at 111 of 154 member institutions, about a 20% response rate and 84% from very-high-research universities, puts 44.3% of federally funded research time on administrative and compliance tasks, up from 42.3% in both 2005 and 2012: proposal preparation 16.0%, post-award administration 13.7%, report preparation 8.3%, pre-award administration 6.3%, active research 55.7%. The strongest single predictor of time away from research was the number of proposals submitted, and researchers submitting twenty or more spent 54.7% of time on requirements. Frontier A contest model puts the theoretical bound: at low paylines the effort wasted writing proposals may be comparable to the total scientific value of the research the funding supports, and as paylines approach single awards the waste and the value converge on zero net benefit. Its structural result is the one to carry: under a threshold-plus-lottery design, dissipated effort depends on the qualifying line rather than the payline, which decouples waste from scarcity. Established A survey of the cost literature assembles the rest — 25–50 days per application against success rates of 10–25%, therefore 100 to 500 person-days of effort per funded project; a European programme evaluation finding 30% to 50% of the funding spent on grant writing; and one programme with a success rate below 1%, six of 650 applicants — and concludes exactly where this brief opens: we do not know what the best distribution system is. Established The obvious remedy was tested once at national scale and failed: the Australian national health funder cut form fields from 180 to 68 and applications from about 100 pages to about 50, and total researcher time rose from 547 to 614 working years while success stayed at 21%. That result belongs to the governance brief and is taken as given here; this brief adds only its funding-design reading — effort tracks the expected value of the prize and the level of competition, so burden is a property of the payline and the award size, not of the paperwork.

Established The state’s largest instrument for shaping what gets built is not its research budget, and the ratio is the first number this subject owes. Public procurement across the OECD runs at roughly 12% of GDP and about a third of general government expenditure, while government-financed research and development in the same economies runs well under 1% of GDP. Frontier A government that wants to change what gets developed therefore holds a purchasing instrument an order of magnitude larger than its granting instrument, and almost nothing in the funding-design literature above is about it. Established The distinction that matters is contractual rather than budgetary: a grant buys an attempt and is terminated rarely; a procurement contract buys a delivered thing and is terminated at a gate. Those are different failure regimes, and the whole procurement argument is about which failures the buyer is willing to own. Handwave Totalling the two — which the market-shaping advocacy routinely does — ignores that most of the twelve per cent buys stationery.

Frontier The milestone contract has the one thing this brief says the whole field lacks: a published cost comparison against the mechanism it replaced. Under its commercial cargo programme an American space agency paid fixed sums on completion of defined technical milestones, under agreements that are deliberately not conventional procurement contracts, and then published its own estimate that developing the resulting launch vehicle under its traditional cost-plus approach would have cost roughly $3.6 billion against an actual development cost near $0.4 billion. Frontier Mark the source and the method: the buyer estimating what it would itself have paid, with a parametric cost model calibrated on its own historical programmes, no control group and no counterfactual vehicle. It is a model output rather than a measurement, and it is still the closest thing to a like-for-like mechanism comparison anywhere in this brief. Established The successor crewed programme transferred development risk the same way at fixed price, and one of its two providers absorbed overruns in the high hundreds of millions against its own accounts rather than the buyer’s. That is what risk transfer looks like when it works, and it requires a supplier solvent enough to survive it.

Established The best-identified evidence in the procurement corner is a regression discontinuity, and it is better identified than most of what is above it. Howell exploits the scoring cutoff in an American energy department’s small-business research awards and finds that an early-stage award roughly doubles the probability that a firm subsequently raises venture capital, while raising patenting and revenue, with effects concentrated in young small firms and absent for the larger later-stage award. Frontier The mechanism the paper identifies is financing rather than certification: the award relieves a capital constraint at a stage where private finance will not price the risk. Established This is the cleanest case in the brief of a funding instrument compared against a defined counterfactual, and it exists because the programme scores applications and cuts at a line, which makes the comparison available without anyone deciding to publish it. Frontier The design lesson generalises past the programme: a funder that already ranks applicants has an evaluation built into its own paperwork, and the reason the rest of this brief has no comparisons is that most funders neither rank on a continuous score nor keep the near-miss.

3 · Frontier questions

Frontier The organising question is whether funding design steers topics or steers risk, and the best evidence says the two levers have very different prices. On topics, work using targeted funding opportunities finds that the switching costs of science are large — inducing researchers to work on marginally different topics requires substantial money in expectation — while also finding that the extra cost may be more than offset by more productive scientists being drawn to those grants. Frontier The specific elasticities are in the published article and are not restated here, because the research pass behind this brief obtained the abstract rather than the full text. Frontier Read against the investigator comparison, this is the sharpest statement of the priority-setting problem available: a funder can change what a scientist works on, at a high and rising price; a funder can change what a scientist attempts comparatively cheaply, by changing the contract term and the failure tolerance. Speculative If that asymmetry is real, funder priority-setting is a weaker instrument for steering topics than for steering risk, and most public argument about research priorities is aimed at the expensive lever.

Frontier Second: is the payline the binding constraint, with every other reform downstream? The contest model says dissipated effort converges on the value of the research at low paylines, and the Australian streamlining result says shortening the form raised total time. Frontier Third, and following from it: threshold-plus-lottery dominates ranking on cost grounds alone, regardless of outcomes, because dissipated effort under that design depends on the qualifying line rather than the payline. That is the strongest live case for randomisation and it requires no outcome comparison at all; it would be falsified by evidence that panels discriminate usefully inside the fundable pool, which is precisely the question the governance brief documents as unsettled.

Frontier Fourth and fifth are two positions held by the same authors about the same model, and both deserve stating. The ARPA model works and its features are complements — flexibility, bottom-up design, discretionary selection and active management raising each other's returns, so partial adoption underperforms. It would be settled by an agency randomising active management across matched portfolios, which the authors of the only measurement say has not happened. The ARPA model is unevaluable in principle on political timescales — the same authors' caveat, and arguably the more important claim. Speculative If the second is true, every replication is a bet on a mechanism that cannot be scored before its authorisation expires, and the growing international cohort of ARPA-style agencies is a set of bets nobody will be able to settle.

Speculative Sixth: focused research organisations fill a real gap that neither grants nor venture capital reaches. Held by the model's promoters, with about a dozen launches, two named technical results and strong reported recruitment as evidence, and no comparison of any kind. Speculative Seventh: prizes are certification devices rather than payments. Supported by medals outperforming money, prize cash covering about a third of invention value, and the online-innovation finding that career and social motivations outweighed cash. It would be settled by a modern prize randomising the medal-versus-cash split, which is cheap and has not been done.

Speculative Eighth and ninth are the two defences of concentration and they are different arguments. Concentration is efficient because talent is concentrated — the standard defence, against which stand the diminishing-returns curve and the finding that within-group inequality exceeds between-group inequality on every grouping tested, and for which stands nothing published identifying the counterfactual productivity of the marginal defunded investigator. Concentration is a self-fulfilling artefact of track-record conditioning — if allocation scores carry information mostly about who the applicant already is, as the governance brief's specification comparison suggests, then concentration is what the mechanism is for rather than a defect in it.

Speculative Tenth: people-not-projects beats project funding inside a public agency. The mechanism exists at scale — 6,053 single-award grants across the agency in 2019–2021 against 224,582 projects, under 3% of grants, of which 4,807 came from one institute — and in 2021 that institute's funding rate for the mechanism ran around 60% against roughly 20% for established-investigator project grants, with higher average funding and recipients submitting fewer applications elsewhere. The entire outcome evidence base is unpublished data presented in a director's slide in May 2022, showing slightly higher productivity and citations and a more diverse grantee pool, reported by an advocacy organisation arguing for the mechanism's expansion. Handwave One unpublished slide is the evidence for a design that reformers propose to scale across a national funder.

Speculative Eleventh, twelfth and thirteenth are the positions that would make this whole brief beside the point, and they are unrefuted. Baseline funding for everyone above a low bar is a live minority European position whose stated obstacle is definitional — who counts as a researcher — rather than evidential. Funding-model choice barely matters because the binding constraint is the supply of good ideas and people is the strongest deflationary claim available: it is consistent with the small size of most measured effects and inconsistent with the top-1% result, and if true, every model here is a redistribution mechanism with a different overhead and the correct optimisation target is overhead. Established And every outcome measure available is an output of the machinery being evaluated — publications, citations, patents and follow-on funding are all produced by institutions inside the system. Handwave That circularity is established; that anyone has proposed a way out of it is not.

Established The philanthropic layer belongs here as a funding model and is now large enough to change the arithmetic. Sector indicators put 2024 philanthropic support for basic and applied research at universities and non-profit research organisations at $27.0 billion$18.3 billion current giving plus $8.8 billion legacy — against $63.6 billion federal, making philanthropy 21% of the total and the federal government 50%, up from $26.2 billion in 2023. Frontier Mark the interest: the body publishing those indicators exists to grow science philanthropy and publishes projections of federal decline alongside them. Frontier Whether such a funder is a legitimate, accountable or well-governed institution is a different question and is not argued here; this brief treats a private funder exactly as it treats a public one — by its contract terms, its horizon, its selection rule and its termination rights.

Frontier Fourteenth: the binding constraint on market shaping is the competence of the buyer rather than the design of the instrument. The mission-oriented literature argues that a public buyer needs the capability to specify an outcome rather than a product, to run a technical dialogue with suppliers, and to terminate a contract without institutional embarrassment — capabilities that three decades of contracting-out reform deliberately moved outside the state. Speculative The claim is plausible, its advocates are the same people proposing the missions, and no published study measures buyer capability and procurement innovation outcomes on the same units, which is what would settle it. Frontier Fifteenth: demand aggregation is the part of the pull toolkit that has actually scaled — pooled purchasing across many small buyers turning a fragmented market into one order large enough to justify a production line — and aggregation rather than the premium is what most of the vaccine record is evidence for. Speculative If that reading holds, the transferable instrument for a mid-sized national buyer is a purchasing consortium rather than an advance market commitment, and the two are constantly conflated.

4 · Technological bottlenecks

Established The first bottleneck is that every outcome measure in this subject is produced by the machinery being evaluated. Publications, citations, patents and follow-on funding are all generated by journals, offices and investors inside the system whose allocation rule is under test. Handwave That is not a reason to stop measuring; it is a reason to distrust any comparison that turns on a small difference in one of them, and nobody has proposed a measure outside the loop.

Established The second is the evaluation window, and it is a bottleneck of politics rather than of method. Projects run about three years; the industry transformations that justify a high-variance funding model take decades; and the model's own advocates state that it is impossible to measure the incidence of one-in-a-thousand ideas on a timescale relevant to programme authorisation. Frontier Any agency evaluated on its authorisation cycle is evaluated on intermediate outputs — publications, patents, follow-on investment — which are exactly the measures the previous paragraph says to distrust.

Frontier The third is that the one well-identified comparison rests on 69 treated scientists. Propensity weighting against NIH early-career prize winners handles selection on observables; selection on unobservables is handled by argument. Speculative The result is strong, replicated nowhere, thirty years old in its treatment cohort, and confined to the life sciences. It is the best evidence in this subject and it is one study.

Established The fourth is that measured discretion is not measured management. The ARPA-E study documents how much programme managers actually change — timelines in 85% of projects, budgets in 43%, milestones in 45% — and its authors state explicitly that because those changes are not applied randomly they cannot claim an impact of active management on performance. Frontier Every internal comparison is inside an agency that manages actively everywhere, so the available contrast is degree, not presence, and the model's central feature has never been compared against its absence.

Frontier The fifth is that application burden is endogenous to the design. The obvious fix was tried at national scale and burden rose: form fields cut from 180 to 68, page counts roughly halved, and total researcher time up from 547 to 614 working years at an unchanged 21% success rate. Established If effort tracks the expected value of the prize and the level of competition, then no administrative simplification can reduce it, and the only levers that can are the payline, the award size and the shape of the contest.

Frontier The sixth bottleneck belongs to procurement and is a rule nobody publishes: how a programme stops. A milestone contract’s advantage over a grant is that it has gates, and a gate is only real if the buyer will walk through it — yet no agency in this brief publishes a termination rate, a distribution of milestones missed before cancellation, or what became of the terminated teams. Established The nearest available figure is internal rather than terminal: 85% of ARPA-E projects had timeline changes and 45% had milestones added or deleted, which is evidence of renegotiation rather than of enforcement. Handwave A mechanism sold on its failure tolerance, whose failure statistics are unpublished everywhere, is being argued for from its design documents. Frontier Termination rates sit in the same administrative records as everything else this brief says is withheld, and they would be the cheapest single disclosure in the subject.

5 · Research dependencies

Established This subject waits on one thing more than any other, and it is not a discovery: a published outcome comparison between an alternative funding model and the mechanism it replaces. The governance brief documents the best-known instance — thirteen years after the first randomised funding round, no funder anywhere has published an outcome comparison between lottery-allocated and panel-allocated grants, despite assignment above the quality threshold being random by construction. Established This brief's contribution is that the lottery case is not an exception; it is the general condition of the funding-model layer. The same absence holds for ARPA-style agencies, for focused research organisations, for prizes, for people-not-projects awards and for block funding.

Frontier The reason is not that the research is hard. In every case the treatment assignment and the outcome data sit in the funder's own administrative records. What is missing is a decision to publish. Frontier That is the same diagnosis the governance brief reaches about integrity data and protocol deposit, arriving independently one layer up: the institutions that could settle these questions hold the data and have chosen not to publish the comparison, and the people best placed to close the gap are the people whose programmes would be scored.

Frontier Second, an evaluation window longer than a programme authorisation. This is a budgeting and statutory dependency rather than a methodological one. A high-variance model whose justification is a small number of very large successes cannot be scored on three-year outputs, and every existing assessment of such an agency has been. Speculative A standing outcome-tracking obligation attached to a cohort of awards, surviving the programme that made them, is the institutional object that does not exist.

Frontier Third, a funder willing to vary contract length and failure tolerance experimentally while holding topic solicitation constant. That is the design that would separate the two levers — risk and topic — whose prices appear to differ by an order of magnitude. Speculative It requires a funder to admit that it does not know which of its own design parameters is doing the work, which is an institutional obstacle rather than a scientific one.

6 · Required experiments

Established The cheapest and highest-value experiment in this subject has already been run and merely needs reporting. Wherever a funder has allocated part of a portfolio by a rule that assigns quasi-randomly above a quality threshold, the comparison exists in its records. The governance brief records the lottery case in detail and finds nothing published. Frontier The same is true for milestone-terminated against non-terminated portfolios, for people-not-projects awards against project grants at the same institute, and for promoted-against-scored selections inside an ARPA-style agency — where 49% of selected applications fell below a hypothetical peer-review cutoff, which is a treatment assignment with a control group attached.

Frontier Second, randomise active management across matched portfolios. The measurement study documents the practice and disclaims the causal reading; the design that would supply it is an agency assigning a subset of projects to a light-touch management arm. Speculative The obstacle is that an agency whose identity rests on active management would be running an arm in which it deliberately manages less, and the political cost of that arm producing better outcomes is obvious.

Frontier Third, randomise the medal-versus-cash split in a modern prize. The strongest historical result in the prize literature is that certification outperformed money, on a nineteenth-century agricultural society. Speculative A contemporary prize could vary the ratio of recognition to cash across categories at near-zero cost and settle whether the certification finding generalises — and given that over 40% of prizes are never evaluated at all, the marginal value of a single well-designed one is high.

Frontier Fourth, vary the payline exogenously and measure dissipated effort. The contest model predicts that waste converges on the value of the research as the payline falls, and the streamlining natural experiment is consistent with effort tracking competition rather than paperwork. Speculative A funder splitting a call into two arms with different announced success rates and measuring preparation time would test the model's central prediction directly, and preparation time is already routinely surveyed.

Speculative Fifth, the one design that would separate risk from topic: hold the solicitation fixed and vary the contract — term length, renewal criterion, tolerance of a failed aim — across otherwise identical awards. Handwave Everything in the investigator comparison points at the contract as the operative variable, and nothing in the record isolates it from the institution, the field or the cohort.

7 · Engineering requirements

Established A funding model is a contract, and the contracts differ on four parameters that can be read off the published terms. Horizon: seven years renewable indefinitely on scientific review at one private institute; eight years for a core investigator at a standing institute with no requirement to seek external grants; three to seven years for a focused research organisation; typically three to five for a project grant. Size: about $11 million per investigator across a seven-year term; $20–50 million for a focused research organisation of 10–30 staff; $1 million over five years and $100,000 over one for the two smaller instruments at the standing institute. Selection: panel score, programme-manager discretion, or invitation. Termination: renewal review, milestone failure, or nothing before the term ends.

Established The programme-manager model's engineering is the term limit and the discretion, and both are documented. About 100 programme managers for roughly $3 billion of extramural funding, each serving a single three-to-five-year term before returning outside; 85% of projects with timeline changes, 43% with budget modifications, 45% with milestone additions or deletions, and 70% experiencing at least one programme-director transition; and roughly 49% of selected applications promoted past a hypothetical score cutoff. Frontier The turnover result is the surprise: projects that changed programme director were 15% more likely to file patent applications, against the authors' own hypothesis that turnover would hurt.

Frontier The milestone contract is the feature that has been tested by transplant and failed. The focused-research-organisation model borrowed it, and the incubator's own account reports that rigid milestone-setting failed and was replaced with adaptive intermediate goals and dynamic quarterly objectives, with product-development demands underestimated. Speculative Whether that is a fault of the transplant or of the original is unresolved, and it is the only against-interest evidence anyone has published about the model.

Established Pull mechanisms are engineered on a different axis: they specify the output and leave the method open. The advance market commitment specified doses and a price; it delivered more than 150 million children immunised against a technologically close target and its designers say a valid counterfactual is lacking. The ten-year prize specified a diagnostic performance threshold and delivered a system doing it in about 15 and 45 minutes. Frontier The design lesson is that pull works where the specification can be written, and the specification can be written where the science is nearly done — which is exactly the domain where pull is least needed.

Established And the system's fixed cost is a design parameter that nobody sets deliberately. 44.3% of federally funded research time on administrative and compliance tasks, 16.0% of it on proposal preparation alone, 25–50 days per application against 10–25% success rates, and therefore 100 to 500 person-days per funded project. Frontier Under a threshold-plus-lottery contest, dissipated effort depends on the qualifying line rather than the payline, which is the only structural proposal in the record that reduces the cost without changing the money.

8 · Adjacent technologies

Within this map the immediate neighbour is Scientific Governance Models, which owns everything that happens to a proposal or a paper once it is inside a process — peer review as a scoring rule, blinding, registered reports, replication, integrity enforcement, and lotteries as an allocation mechanism. This brief uses its lottery finding and its streamlining result and re-derives neither. Also Scientific Advisory Institutions, which owns the machinery by which scientific priorities reach government; Artificial Scientists, which inherits both the funding contract and the literature quality floor; Long-Term Institutions, where an endowed research institute is one of the few working long-horizon designs; Institutional Design, where the general problem of an agency evaluated on the wrong horizon is stated abstractly; and Future Public Administration, whose finding that evaluation is discretionary explains most of the absences recorded here.

Private giving as an institution — its tax treatment, payout, accountability and the legitimacy critique — is a different subject, and this brief takes none of it. A private funder is treated here exactly as a public one: by horizon, selection, contract and termination.

Outside it: the economics of innovation, which supplies the prize and advance-market-commitment evidence; contest theory, which supplies the payline result; the sociology of science, which supplies the novelty measures; and public-sector budgeting, which supplies the authorisation cycle that determines what can ever be evaluated.

9 · Institutional requirements

Established The institutional finding is a publication decision, not a knowledge gap, and it is the single most important claim in this brief. Every alternative model here is operating with real money: $3.5 billion across 1,465 projects at one agency, more than 250 investigators on renewable eight-figure contracts at one institute, about a dozen focused research organisations mid-flight, $650 million committed to a no-grant institute, roughly 5,000 people-not-projects awards from a single institute, a completed £8 million ten-year prize, and partial randomisation at more than a dozen funders. Established For exactly one — the privately funded investigator programme — is there a published comparison against the mechanism it replaces, and it was produced by outside economists rather than by the funder.

Established What exists instead, itemised, because the list is the argument. Cumulative output counts published by the agency itself. A within-agency correlational study whose authors explicitly disclaim causal interpretation. Two named technical results in an essay written by the incubator's own staff. One unpublished slide. An operator's announcement listing achievements and stating no limitations. And, for lotteries, the documented nothing the governance brief records. Frontier Advocacy for funding-model reform has run more than a decade ahead of the evidence for it.

Established Interested parties are everywhere in this subject and the directions differ. The agency reports its own impact and calls the figures key early indicators. The incubator's staff write the fullest public account of the model they incubate — and also publish that its milestone feature failed, which is the most valuable sentence in the source. The prize operator publishes that prize evidence is scarce and that over 40% of prizes are never evaluated. The advocacy organisation reporting the people-not-projects productivity comparison wants the mechanism expanded, and is reporting unpublished funder data. Established Two findings run hard against their publishers' interest — agency leadership documenting concentration in their own distribution, and the advance market commitment's own designers stating that they lack a valid counterfactual — and both are weighted up for it.

Frontier And the structural reason the comparisons are not published is visible in the incentives. The funder holds the assignment and the outcomes; the funder chose the design; and a published comparison can only confirm what the funder already believes or embarrass it. Speculative Nothing in the record suggests suppression. Everything in the record is consistent with nobody's job description including it, which is the more common and more durable failure mode.

10 · Ethical & societal considerations

Established The distributional evidence is unusually good and it points somewhere unexpected. Within-group inequality exceeds between-group inequality on every grouping tested — career stage, gender, race, degree, organisation, region, state. Frontier A reform aimed entirely at closing between-group gaps is aimed at the smaller component of the problem, which does not make it wrong and does mean the arithmetic should be stated. The same authors, publishing about their own agency, call the racial and gender components unacceptable inequalities.

Frontier The competitive-funding result carries a distributional finding that its own headline hides. Competitive funding raised measured novelty on average, and the effect reverses for junior researchers and for female researchers — interaction coefficients of −0.94 and −0.95. Speculative If that holds outside the one country it was measured in, competition is not a neutral instrument: it rewards the already-established and penalises the same characteristics that the between-group reform agenda is aimed at, through a mechanism nobody designed.

Established The application system imposes a cost that falls unevenly and is measured. 44.3% of federally funded research time goes to administration and compliance, rising to 54.7% for researchers submitting twenty or more proposals in three years. Frontier That cost is borne disproportionately by people who must submit more to get funded, which is the people with weaker track records — a regressive burden generated by a mechanism whose stated purpose is merit selection.

Speculative And the high-variance contract has an ethical shape that is rarely stated. The investigator comparison shows the design buying more top-percentile output and more work below the scientist's own prior floor. Frontier A funder choosing that contract is choosing to purchase failures deliberately, on the argument that the tail justifies them. That is defensible and it is a choice about other people's careers, made by a funder who bears none of the reputational cost of the failures.

11 · Civilizational implications

Frontier The civilizational claim this brief can support is narrow and worth stating precisely: a funder can change the risk profile of a research portfolio, and there is one measured demonstration of it. A long renewable term with review on scientific merit rather than delivery produced +97% in top-percentile output and +35% below the investigator's own prior floor. Speculative If that generalises, the design lever that matters for the long-run rate of scientific surprise is contract structure rather than money, and it is cheap.

Speculative Against that stands the deflationary position, and this brief cannot refute it: funding-model choice may barely matter, because the binding constraint is the supply of good ideas and people. It is consistent with how small most measured effects in this subject are. It is inconsistent with the top-1% result. Handwave If it is right, every model here is a redistribution mechanism with a different overhead, and the whole reform argument should be about overhead — which, given that 100 to 500 person-days go into every funded project, would still be a large prize.

Established The most durable observation is about institutions rather than about science. A dozen serious alternatives to project-based peer review are running simultaneously with billions of dollars behind them, and the comparison that would tell anyone which of them works sits unpublished in administrative records. Frontier A civilisation that funds science this way is not choosing between designs; it is running them all and declining to score them.

12 · Timelines

These horizons track programme authorisations, contract terms and publication decisions rather than technology:

  • 10 yr: Frontier The first cohort of focused research organisations reaches the end of its three-to-seven-year terms inside this window, which is the first opportunity anyone has had to count completed cases. Established The British randomisation trial states it will examine award outcomes during 2025–2028, which is the nearest thing to a scheduled outcome comparison anywhere in the record. Speculative Expect ARPA-style agencies to continue proliferating internationally and expect none of them to publish a counterfactual, because the authorisation cycle rewards output counts. Frontier Expect procurement-side instruments — milestone contracts, purchasing consortia, advance commitments — to be adopted faster than granting reforms inside this window, because a buyer can change a contract without new authorising legislation, and expect their termination statistics to stay unpublished.
  • 25 yr: Speculative Either standing outcome-tracking obligations outlive the programmes that make awards — in which case the high-variance models become evaluable for the first time — or they do not and the field argues from cumulative output counts indefinitely. Speculative The investigator comparison's treatment cohort will be sixty years old by then and still the only well-identified funder comparison unless someone runs another. Handwave Nothing in the current incentives suggests anyone will.
  • 50 yr: Speculative If the topic-versus-risk asymmetry is real, national research strategies converge on setting contract terms and failure tolerance rather than on naming priority areas, because the second lever is an order of magnitude more expensive. Speculative If the deflationary position is right instead, the whole apparatus converges on minimising overhead, and the historical record of this period reads as an expensive exploration of designs that were roughly equivalent.
  • 100 / 250+ yr: Handwave Beyond useful forecasting. The one observation with that reach is that the endowed institute with a long horizon and no requirement to reapply is an old form that predates competitive grantmaking, and that the modern experiments in it are re-inventions rather than inventions — which is a remark about institutional memory, not a base rate.

13 · Technology tree & dependencies

  • Depends on Nothing on this map. This brief waits on no result another brief produces. It uses Scientific Governance Models' finding on lottery allocation and its national streamlining result as established inputs, and adds the funding-model layer above them; that is a relationship of inheritance rather than of dependency, and no typed depends-on edge is claimed for it.
  • Requires (not on this map) Four institutional decisions, none of them a research result, and all three currently withheld. First and centrally, a published outcome comparison between an alternative funding model and the mechanism it replaces. Every alternative on the reformer's list is running with real money — $3,508.2m across 1,465 projects at one agency, more than 250 investigators on renewable seven-year contracts worth about $11m each, about a dozen focused research organisations mid-flight, $650m committed to a no-grant institute, 6,053 people-not-projects awards across a national funder in 2019–2021 with 4,807 from one institute, a completed ten-year £8m prize, and partial randomisation at more than a dozen funders — and for exactly one, the privately funded investigator programme, has anyone published a comparison, produced by outside economists rather than by the funder. FR-VIII-06 documents the best-known instance: thirteen years after the first randomised round, no funder anywhere has published an outcome comparison between lottery- and panel-allocated grants. That absence generalises. In every case the assignment and the outcomes sit in the funder's own administrative records, so what is missing is a decision to publish. Second, an evaluation window longer than a programme authorisation: projects run about three years, the transformations that justify a high-variance model take decades, and the model's own advocates state that the incidence of one-in-a-thousand ideas cannot be measured on a timescale relevant to political decisions about authorisation — so every assessment scores intermediate outputs produced by the machinery being evaluated. Third, a funder varying contract term and failure tolerance experimentally while holding topic solicitation constant, which is the only design that would separate the two levers whose prices appear to differ by an order of magnitude: topic switching is expensive, and changing what a scientist attempts is not. Fourth, a published termination rate for milestone-funded programmes: the procurement instruments now spreading through research funding — milestone contracts, staged small-business awards, advance commitments — are sold on the proposition that they stop paying when a team misses a gate, and no agency publishes how often it stops, at which milestone, or what happened next. The one adjacent figure in the record is renegotiation rather than enforcement: 85% of projects at one agency had timeline changes and 45% had milestones added or deleted. All four are things a funder or a legislature could choose to supply and have so far chosen not to.
  • Enables In principle every research programme on this map is allocated by one of the mechanisms described here, and the contract terms recorded above determine what any of them can attempt. No typed enabling edge is claimed, because the relationship has never been measured: for no funding model in this brief does a published comparison exist against the mechanism it replaces, so the effect of a funder's design on any downstream research outcome is unquantified in every direction.
  • Adjacent The economics of innovation, which supplies the prize and advance-market-commitment record; contest theory, which supplies the payline result; the sociology of science, which supplies the novelty measures; public-sector budgeting, which supplies the authorisation cycle; and within this map Scientific Governance Models, Scientific Advisory Institutions and Long-Term Institutions.

14 · Common misconceptions & speculative claims

Established “The alternatives to grant peer review are known quantities.” They are known as designs and unknown as outcomes. Not lotteries, where the governance brief records that no comparison exists thirteen years on. Not ARPA-style agencies, where the only measurement study disclaims causal interpretation and the authors say the horizon exceeds the window. Not focused research organisations, which have no outcomes at all. Not prizes, where over 40% are never evaluated. Not block funding, which has one cross-sectional survey whose primary hypothesis failed. Not people-not-projects, whose evidence is one unpublished slide. Frontier One exception exists and it shows a different distribution, not a better one.

Frontier “The investigator result licenses copying the model.” The comparison is against NIH early-career prize winners in the 1990s, on 69 treated scientists, in one field, with selection on unobservables handled by propensity weighting rather than randomisation. Established It also shows 35% more output below the investigator's own prior citation floor and 3.2 against 5.1 applications to the public funder, with worse scores on the ones filed. The design bought more failures and disengagement from the comparison institution, and both belong in any summary of it.

Established “The agency's cumulative figures demonstrate impact.” $11.8 billion of follow-on funding, 150 companies and 27 exits at $21.9 billion of market valuation are outputs reported by the agency itself with no comparison group, described in its own report as key early indicators. Speculative The one thing resembling a control — a higher patenting rate than similar awards from other branches of the same department — reaches this brief at second hand through a review rather than from the primary study, and is flagged accordingly.

Established “Prizes reliably induce invention.” The strongest evidence is a nineteenth-century agricultural society in which medals outperformed money and prize cash covered about a third of an invention's sale price. Frontier The largest modern pull mechanism, by its designers' own account, worked on a technologically close target — vaccines already in late-stage trials — and its designers state that they lack a valid counterfactual. A pull mechanism that scaled manufacturing is on the record; one that induced an invention is not.

Frontier “Research funding concentration is worsening rapidly and is worse than in the wider economy.” The top centile of principal investigators went from 8.3% to 10.8% of funding between 1995 and 2019 while the top centile of the population went from 14.3% to 18.7% of income — a 30% relative rise against a 31% one. Established Institutional concentration is extreme and roughly flat in share: the top 100 institutions took 78.4% in FY2023, and the top 10's share fell from 22.7% to 20.1% over a decade while their dollars rose sharply.

Speculative “Capping the largest grants would fund thousands more productive investigators.” The proposal computes about $4.22 billion freed and roughly 10,500 additional investigators from a diminishing-returns curve with a sweet spot near $400,000. Frontier That curve is drawn across investigators the current system selected. Applying it to investigators it rejected assumes the marginal addition produces at the sweet-spot rate, which is an extrapolation rather than an estimate, and it is the load-bearing step of the argument.

Established “Cutting the paperwork cuts the burden.” Tested once, at national scale: form fields from 180 to 68, applications from about 100 pages to about 50, and total researcher time up from 547 to 614 working years at an unchanged 21% success rate. Frontier Effort tracks the expected value of the prize and the level of competition, which makes application burden a funding-design problem rather than an administrative one — and the only structural fix in the record is a contest whose dissipated effort depends on a qualifying line rather than on the payline.

Frontier “Block funding protects adventurous research” and “competition drives it out” are both unsupported. The one study found competitive funding raised measured novelty by +0.35 to +0.60 on average and lowered it for junior and female researchers. Handwave Neither headline claim survives, and both continue to be made confidently by people who have not looked for the evidence, because the evidence is one survey in one country whose authors predicted the opposite of what they found.