1 · Concept overview
AI-assisted governance means the state using algorithmic systems to administer: to decide who is eligible for a payment, which claims to investigate, which debts to raise, how to triage an application, where to send an inspector, and lately to draft the submission on which a minister decides. It is the machinery of public administration rather than the machinery of policy, and the people it acts on are usually claimants rather than customers — they cannot take their business elsewhere.
The question this brief owns is narrower than its title and is the one least often asked: does any of it make governing better? Not whether it is lawful, not whether it is fair, not what it does to the people who work in the building — whether decisions become more accurate, services reach more people faster, or policy analysis improves. That question now has an evaluation literature. The literature is small, its best-designed members disagree with its most-quoted members in a systematic direction, and not one member of it measures output.
This is not the same subject as AI Governance, and the two titles are close enough that a reader will assume otherwise. That brief is about how a society governs AI — the rules a polity writes for a technology. This one is about how a government administers using algorithms. The collision check that scoped this category scored the two at zero term overlap, which was correct on the terms and is exactly why the boundary has to be stated here in prose: nothing in the machine-readable layer will carry it.
Two sibling boundaries matter more, because both slots hold material this brief also uses. Digital Constitutional Systems owns the constitutional and rights questions about automated decisions: judicial review of coded administration, what a person is owed by way of explanation, rights entrenchment, and whether machine-readable rules can be law. Where this brief cites the same judgments — SyRI, Robodebt, the Dutch regulator’s fines — it takes from them only what they establish about how a deployed system operated, and leaves the constitutional analysis there. Future Civil Services owns the workforce: headcount, pay, tenure, skills, and whether machines substitute for administrative labour. The three United Kingdom departmental evaluations of Microsoft 365 Copilot appear in both briefs and are read for different things — there, whether the service needs fewer people; here, whether the government’s output improved. Neither question can be answered from the other’s data, and the fact that the same three studies are the only evidence for either is itself the finding.
One further scoping decision shapes everything below. This brief is organised around deployed systems with documented outcomes, and it applies an explicit evidence hierarchy: randomised or directly observed measurement first, then comparison-group designs, then self-report, then press releases and vendor material. That hierarchy does most of the analytical work in the sections that follow, because in this subject the reported size of the benefit falls monotonically as you climb it.
2 · Current scientific position
Established The central finding of this brief is a dose–response relationship between study design and reported benefit: the better the design, the smaller the effect, and the best-designed study found users slower and less accurate. Three United Kingdom departmental evaluations of the same product — Microsoft 365 Copilot, in overlapping periods across 2024–25 — ran under three different designs and produced three different answers, and the answers are ordered by rigour.
Established The widely quoted figure is 26 minutes saved per day, annualised to thirteen days a year. It comes from a cross-government experiment covering 20,000 users in twelve organisations, with survey responses from 7,115 and telemetry from 14,500. It is entirely self-reported and there was no control group. The report’s own limitations section says the savings “are self reported” and should be read alongside literature on what such metrics are worth, and that “it was not possible to identify how time saved was spent.” Seventeen per cent of respondents reported no clear saving at all.
Established The Department for Work and Pensions re-ran the exercise across 3,549 licences with a comparison group — 1,716 Copilot users against 2,535 non-users, adjusted for demographics, role and prior AI experience — and the figure fell to 19 minutes a day. Participants were volunteered or nominated rather than randomised. Simply introducing a comparison group cut the effect by roughly a quarter.
Established The Department for Business and Trade ran 1,000 licences, about 70% volunteers and 30% randomly selected and stratified by directorate and grade, and did the thing nobody else did: it recorded eleven tasks on video and had them blind-scored from one to five for quality and accuracy. Its conclusion, in full: “The evaluation did not find evidence that time savings have led to improved productivity, and control group participants had not observed productivity improvements from colleagues taking part in the M365 Copilot pilot.” On spreadsheet data analysis, users of the tool were slower — 25:01 against 20:33 — and less accurate, scoring 1.5 against 2.7. Observed-task samples were three per condition and the report says so.
Handwave The 95-minute figure that circulates in North American coverage sits below all three. It is an exit survey of 136 respondents out of 175 volunteers across fourteen agencies, in a sample the report itself calls unrepresentative and whose limits it sets out candidly. The misconceptions section takes it in full; what matters here is where it lands on the ladder.
Established And underneath all three sits the omission that disciplines this entire brief: no published study of an AI deployment in a government agency measures output. The cross-government report could not identify how the saved time was spent. The Business and Trade report states that productivity “was not a key aim of the evaluation.” The Work and Pensions study measured minutes. On the evidence assembled here there is no published evaluation of AI in a government agency that reports cases cleared, decisions made, backlog reduced, error rate changed or staff required. Every headline number in this field is a measure of self-perceived time on task, and the conversion rate from time on task to governing is not merely unfavourable — it is unmeasured.
Established There is a strong prior against AI-as-decision-improver, and it comes from the one part of institutional design with a real engineering record. Market design put a deferred-acceptance clearinghouse into New York City’s high-school assignment in 2003. The effect was large: administrative assignments — students placed at a school none of them had asked for — fell from 26,098 (37% of the cohort) to 7,143 (10.7%) in one year. Then Abdulkadiroğlu, Agarwal and Pathak decomposed the gain in units of willingness to travel. The distance between neighbourhood assignment and the utilitarian optimum is 18.96 miles; the coordinated mechanism closed all but 3.73 miles of it, about 80% of the available welfare. Switching to the student-optimal stable matching would have added 0.11 miles, or 0.6% of the range. A fully Pareto-efficient matching would have added 0.62 miles, or 2.7%. In the authors’ words, algorithmic improvements are “swamped by the effect of simply having choice in a coordinated system.” The clever part of the design did almost nothing. Replacing chaos with any deadline-respecting queue did nearly everything.
Established The same conclusion arrives from the opposite direction in auctions. Klemperer’s comparison of the 2000 European third-generation spectrum auctions found per-capita revenues of €630 in the United Kingdom, €615 in Germany, €240 in Italy and €170 in the Netherlands — a near-fourfold spread within one year driven by how many serious bidders turned up rather than by mechanism, with the one weak Dutch entrant withdrawing after a threatening letter from an incumbent’s lawyer. His conclusion is that outcomes turn on “discouraging collusive, predatory and entry-deterring behaviour” and that design is “horses for courses, not one size fits all.” In schools the clever part of the design was worth almost nothing because coordination did the work; in spectrum it was worth almost nothing because market structure did. In both the mechanism was necessary and the mechanism was not what varied. Anyone claiming that a better model will improve public decisions is claiming the residual these two decompositions measured and found small.
Established The successes in this field with genuine causal evidence are consistent with that prior, because none of them is a prediction. The most robust result in the administrative literature is a randomised Danish tax-audit experiment on more than 40,000 filers: evasion on third-party-reported income is 0.23% against 17.1% on self-reported income. That is the mechanism underneath Nordic tax pre-filling, and it is informational architecture. The second is that automatic pension contributions pass through to total saving at about 78 cents in the krone while tax subsidies raise total saving by roughly one cent per dollar spent, because around 85% of people are passive. The third is outside government entirely: Feeding America replaced needs-based rationing of donated food with a scrip auction in 2005 and fed 55,000 additional people a day in the first year, because prices aggregated information about which food banks could use which loads. Getting the right data from the right party, changing the default, and coordinating a queue. No model appears anywhere in any of them.
Established The strongest randomised test of a large government automation programme found the efficiency gain undetectable and the exclusion cost measurable. Muralidharan, Niehaus and Sukhtankar exploited a randomised rollout of biometric authentication across 15.1 million beneficiaries in 132 blocks of Jharkhand’s public food distribution system. Biometric authentication alone produced no significant reduction in leakage; the confidence interval on the change spans −1.7% to +6.5%, including zero and including an increase. Beneficiary transaction costs rose 17%. For the 23% of households whose records were not correctly linked, benefits received fell 8.4% and the probability of receiving nothing rose 10 percentage points. Under the later reconciliation phase, which did cut disbursals, 22–34% of the reduction came out of beneficiaries’ receipts rather than out of leakage. The authors’ own decision rule: a planner would need to value marginal fiscal revenue at at least 28% of the value placed on transfers to marginal households before the policy was worth adopting. This is not an AI system, and that is the point — it is the highest-quality causal evidence anywhere on whether automating an administrative gate improves administration, and the answer was no.
Established When an audit office has tested a government savings claim from a digital system, most of the saving turned out to be something else. India’s government claimed 21,552 crore rupees saved in 2015-16 from identity-linked cooking-gas subsidy transfer. The Comptroller and Auditor General reported to Parliament that of the 23,316.12 crore fall in subsidy payout, 21,552.28 crore was attributable to the collapse in crude prices and only 1,763.93 crore — about 7.6% — to eliminating duplicate and fake connections. The United Kingdom’s identity-verification programme is the same shape audited from the other end: the Cabinet Office’s own benefits estimate for 2016-17 to 2019-20 was revised from £873 million to £217 million, a 75% cut, and the National Audit Office stated it “has not been able to replicate or validate” even the reduced figure. Two audits, two continents, one result: the saving is smaller than claimed and its causes are not the technology.
Established The failure record, by contrast, is judicial, specific and much wider than the two countries it is usually attributed to — adverse findings against deployed government systems in Australia, the Netherlands, the United States, Poland and Austria, a pending challenge in France, and a 2025 study analysing 71 federal and state court dockets contesting algorithm-based determinations in United States disability, unemployment and nutrition assistance. The constitutional reading of that record belongs to Digital Constitutional Systems. What belongs here is the correction it forces: the paradigm case is miscast. Robodebt was deterministic data-matching plus arithmetic — no model, no training data, no score — and the Royal Commission located the illegality in income averaging and the absence of legal authority to demand the information, attributing part of the operational failure to haste. The Dutch childcare-benefit disaster turns the same way: what destroyed families was an all-or-nothing recovery doctrine upheld across 1,875 higher appeals from 2011 to October 2019 by a court that has since apologised for maintaining it too long. In the two cases that define this field, the machine selected who was investigated and the law did the damage.
Established And there is still no denominator anywhere. The United Kingdom’s Algorithmic Transparency Recording Standard became mandatory across central government in 2024 and held 59 records in May 2025 and 143 in August 2026, against a National Audit Office survey of 87 bodies that found 74 AI use cases already deployed. The United States federal inventory reports 3,611 use cases across 56 agencies — and its own auditor found that of twenty agency inventories examined only five were comprehensive, one agency reporting 375 uses privately against 33 publicly, with the Pentagon and Intelligence Community exempt altogether. The Netherlands’ register holds 1,536 entries and registration there is voluntary. No failure rate can be computed, in either direction. Every claim about how often these systems go wrong — including the optimistic ones — is unsupported.
Established The most complete public-sector assurance architecture yet written is American, dates from April 2025, and its central instrument is a waiver. Office of Management and Budget memorandum M-25-21 collapsed the previous framework’s two protected categories — safety-impacting and rights-impacting — into one class of high-impact AI, defined by whether the output is a principal basis for a decision with a legal, material or similarly significant effect on a person, and attached minimum practices to that class: an impact assessment before deployment, pre-deployment testing in a realistic context, ongoing monitoring after go-live, human oversight and operator training, and a route to timely human review and remedy for an adverse determination. Agencies had until 3 April 2026 to bring systems already running into line or stop using them. The escape hatch is in the same document: a Chief AI Officer may waive one or more of the minimum practices in writing, with waivers centrally inventoried and published at least annually. This is not a prohibition with exceptions but an obligation with a documented opt-out, and the count of opt-outs is the number that says what it is worth.
Established The companion memorandum pushes the same duties into the contract, which is the only place they bind a vendor. M-25-22, issued the same day, governs federal AI acquisition: performance-based terms tied to what the system must actually do, protections stopping a supplier from training on government data, explicit intellectual-property and data-rights allocation, monitoring obligations carried through the contract term, and provisions against vendor lock-in — the condition in which the state cannot change supplier because the model, the tuning data and the integration all belong to somebody else. Frontier What the pair does not supply is the thing this brief keeps asking for. An impact assessment is a document, a pre-deployment test is a procedure, monitoring is a duty, and none of the three is an output measure: a system can satisfy every minimum practice and still publish no figure for cases cleared, decisions overturned or error rate changed. The memoranda regulate how the state acquires and runs the system. Neither requires anyone to find out whether it governs better.
3 · Frontier questions
Frontier The override rate is the single most useful diagnostic in this field and almost nobody publishes it. Where a human is said to review an algorithmic recommendation, the measured rate at which they actually change it ranges from 0.58% in Poland’s jobseeker-profiling system, across 341 labour offices and roughly 1.5 million people, to 77% for over-65s in the United Kingdom’s Universal Credit advances model. Those two numbers describe completely different institutions wearing the same label. Until a system publishes its override rate, “there is a human in the loop” conveys nothing at all.
Frontier The most specific and most useful result in the AI-in-government literature is that the tool is good at exactly one clerical task and bad at another, and that finding is one experiment old. The directly observed tasks in the Business and Trade evaluation decompose as follows: summarising a long report was 3.3 times faster with better accuracy — 12:37 against 41:34, scoring 4.0 against 2.5 — while spreadsheet data analysis was slower and markedly less accurate. Slide production was faster and much worse. Scheduling was a net loss. That predicts real substitution in précis-writing and briefing work and none whatsoever in analysis, and it means an organisation that adopts the tool without knowing which of its tasks are which will get a mixture and measure the average. The result is frontier rather than established because the observed-task samples were three per condition, in one department, on one product version.
Frontier The genuinely open question is the conversion function from time saved to output, and it has never been estimated. Suppose the 19-minute figure is right. Nobody has published what a saved minute in an administrative agency becomes: a case cleared, a longer queue cleared no faster, a meeting, or nothing. The cross-government evaluation is explicit that it could not tell. Until somebody instruments a caseload rather than a diary, the entire policy edifice built on these numbers — national digital-exchange savings targets, transformation programmes, workforce plans — rests on a quantity whose relationship to governing is assumed rather than observed.
Frontier The automation-bias premise is not supported by the best-powered study in this setting, and the result runs the other way. Three pre-registered experiments in the Netherlands found no automation bias among citizens, and among 1,345 civil servants found participants significantly less likely to follow algorithmic advice than identical human-expert advice — 4.8% against 8.6%. The selective-adherence hypothesis also failed its key interaction test. These are vignette experiments rather than field data, and the third study ran shortly after the childcare-benefit scandal, a plausible confound in the direction of algorithm-scepticism. But the honest statement of the evidence is not that human review is rubber-stamping. It is that human oversight is substantively active and substantively unaudited.
Frontier The unasked question is the human counterfactual. United States nutrition assistance is a largely human-adjudicated means-tested programme with a FY2023 national payment error rate of 11.68%, ranging by state from 3.27% to 60.37%. Michigan’s automated system was 93% wrong, which is far worse and rightly a scandal. But a literature that compares deployed algorithms against an implicit baseline of correct human administration is comparing against something that does not exist. Whether the algorithmic error rate is above or below the manual one, and whether it is distributed worse, is asked in almost none of this material.
Frontier A harder version of the same problem: nobody has established that “better governing” can be measured at all in the way this field assumes. Vaccaro compared seven established cross-national state-capacity measures. Their pairwise correlations run from 0.70 to 0.94 and a principal component analysis attributes 86.91% of common variance to a single component, which reads as convergent validity. They do not behave as though they measure one thing: 45 countries diverge by more than 0.40 standardised units between two of them, divergence is systematically largest at intermediate capacity where most policy-relevant countries sit, and three published findings about democracy and state capacity reverse or vanish depending on which index is substituted. If a dependent variable can flip a published result across seven indices that correlate above 0.70, then claims that a technology raised or lowered administrative quality need to say which measure they mean before they can be tested.
Frontier The historical question — does computerising an administration raise its output — is under-researched rather than answered. The one directly relevant empirical study located runs against the sceptical intuition: a 2002 test of the productivity paradox in United States state governments reported that “IT investments by state governments have positive and significant effects on the measure of productivity, gross state product,” with larger returns where a chief information officer structure existed. That is one study, one country, one level of government, twenty-four years old, and only its abstract was obtainable, so the number of states, the years covered, the effect sizes and any employment effect are unknown to this brief. The honest position is that the historical question is open, not that it has been settled in the sceptical direction.
Frontier Nobody has yet observed the accountability apparatus stop anything. There is no documented case in which an entry in a transparency register caused a system to be halted, altered or refused. What has actually stopped systems is litigation, investigative journalism, and pilot results. Whether registers can be made to bite is genuinely open and is treated in the institutional section below.
Frontier “Independently tested” is not yet a comparable statement, because nobody publishes the battery, the threshold or the failure. No deployed benefits, fraud or triage system located for this brief has published its pre-deployment test protocol together with the criterion that would have failed it. Canada’s directive comes closest to a hard independent-test clause — external peer review before launch at the higher impact levels — and even there the published record shows reviews performed and none recorded as having stopped a deployment. A test whose failure condition is unpublished cannot be told apart from a test that cannot be failed, and the distinction is cheap to settle: one agency publishing one protocol, one threshold and one system that did not clear it.
Frontier The incident channel is the missing sensor, and its absence is why the failure record in this field is a journalism record. No mandatory incident-reporting duty for algorithmic administration has been located anywhere; the repositories that exist are voluntary and assembled largely from media coverage rather than from reports by the operators of the systems. That inverts how every other safety-critical public function is instrumented — aviation, medicines and transport each pair a compulsory occurrence report with a denominator of operations, which is exactly why they quote rates and this field cannot. So the observed distribution of algorithmic failure is the distribution of investigative attention, and whether any administration adopts a duty to report is the second most informative thing that could happen here, after publishing an output measure.
4 · Technological bottlenecks
Established The first bottleneck is the arithmetic of remediation at scale, and one document establishes it beyond argument. Having held a royal commission into unlawful automated debt-raising, Australia then confronted a second and larger problem: income apportionment, a separate, non-algorithmic, manual practice applied unlawfully from at least 2003, covering 5.5 million debts, 3 million people and $4.4 billion, with an average debt age of nineteen years. Its own Office of Impact Analysis costed the options in September 2025. Identifying and recalculating all 5.5 million debts lawfully: $1,422 million over ten years, potentially two months per debt. Retrospectively validating them instead: $2.8 million. Validation plus a streamlined resolution scheme: $55.9 million. The last was preferred, and affected people receive up to $600 per debt. A state that has just been publicly humiliated for unlawful automated debts will, three years later, choose to legalise 5.5 million unlawful decisions rather than recalculate them, because on its own published numbers doing it properly costs five hundred times more. Errors made at machine speed must be unwound at human speed, and no administration can afford that.
Established The second is quality-assurance capacity, which is the safeguard the whole apparatus assumes and the one nobody has resourced. Between 23% and 64% of users in the only directly observed evaluation did not check the output or could not compare it, and 22% reported hallucinations. A separate government pilot recorded made-up citations and links and a user reporting that on legal questions the answers “were almost never correct.” The design of every deployed system in this brief assumes a human check somewhere; the measured behaviour of humans given a plausible-looking draft is that a large minority do not perform it. That is not a training problem to be solved by a module. It is what a time-saving tool does when the time saved comes from not reading.
Established The third is that the outcome is not instrumented anywhere, so the systems cannot be managed. No published output measure, no published override rates, no verified inventory, and no audit of the human baseline the systems are compared against. Each of these is cheap. None is a research problem. Their joint absence means that a department cannot tell an effective deployment from an ineffective one, and neither can anybody outside it.
Established The fourth is that the inventories which would supply a denominator are known by their own auditor to be unreliable. Of twenty agency inventories examined, five provided comprehensive information; duplicates, non-AI technology and research projects were improperly included; one agency reported 375 uses to the auditor against 33 on its public list; and the Intelligence Community and the Pentagon are exempt from reporting altogether. A use-case count is a reporting artefact, not a deployment measure, and building oversight on top of it is building on a number nobody can reproduce.
Frontier The fifth is that the binding scarcity is coordination rather than prediction, which means most of the effort is aimed at the small residual. The New York schools decomposition puts algorithmic refinement at 0.6% to 2.7% of the available welfare against roughly 80% for coordination alone. If that generalises even loosely to administrative systems, then an agency that has not yet coordinated its queues, deadlines and data flows will get more from doing so than from any model, and an agency that has coordinated them has already captured most of what is available. This is a claim about where the remaining gains are, not a claim that mechanisms do not matter, and the honest caveat is that the decomposition has been performed once, in one sector, in one city.
Frontier The sixth is review capacity, which is the constraint that converts every other problem into a crisis. Six years into the Dutch childcare-benefit recovery operation: about 69,000 applicants, over 43,000 recognised, every recognised victim finally through integral assessment by December 2025 — and roughly 7,400 appeals against those assessments plus about 9,000 pending damage claims still open. The national Court of Audit found statutory deadlines mostly missed and compensation routes so methodologically inconsistent that awards cannot be compared. Repair has taken longer than the harm and has produced a queue of its own.
Established What is not a bottleneck, stated plainly because the field’s discourse says otherwise. Model capability is not the constraint on any system described in this brief. Neither is data volume, compute, nor the state of the research literature. The constraint is that nobody measures the output, nobody can afford to unwind the errors, and the gain being chased is the part of the problem that mechanism design says is small.
Established In the one place where compliance with an algorithm-audit mandate has been counted, the rate was about five per cent. New York City has required since July 2023 that employers using automated employment decision tools commission an annual independent bias audit and publish the result. Researchers then did the rarely-done thing and looked: across 391 employers examined, 18 had published a bias audit and 13 the required transparency notice. Enforcement is complaint-driven and the statute leaves wide room to declare a tool out of scope. This is private-sector and one city, so it transfers to public-sector assurance as a prior rather than a finding — but it is the only empirical prior available, and it says that a published-audit duty without an inspector and a denominator produces paperwork from the minority who were going to comply anyway.
Frontier An appeal right is a claim on a queue, and nobody has costed the queues the new rights create. The American minimum practices require a human-review and remedy route for high-impact systems; the European high-risk regime, when it bites, adds a right to an explanation of the system’s role in an individual decision and a complaint route to a market surveillance authority. Each is a real improvement on no route at all, and each converts an algorithmic error into a case a human must hear, at the remediation cost priced above. The diagnostic that would show whether the rights are real is the one missing everywhere else in this brief: appeals lodged, proportion upheld, median time to decide. Until an administration publishes an overturn rate for its automated determinations, a statutory right of appeal is unfalsifiable in exactly the way “a human reviews it” is.
5 · Research dependencies
Established Nothing here waits on a research result, and that is the defining feature of this topic’s dependency structure. The models are ordinary, the deployments are commercial products, and nothing on the critical path is a discovery. A brief in a frontier-research corpus that depends on no frontier is unusual and the fact should be stated rather than smoothed over.
Established What it waits on is measurement and administrative capacity, in that order. An output instrument — cases cleared, error rate, backlog age — without which no deployment can be evaluated and none has been. Published override rates, without which human oversight is unfalsifiable. A verified inventory, without which no rate of anything can be computed. Quality-assurance capacity at the point of use, since between a quarter and two-thirds of users of the only directly observed deployment did not check the output. A review path with the capacity to reverse decisions at the rate the system generates them. And a funded liability for being wrong at scale, which the Australian remediation arithmetic prices at roughly five hundred times the cost of legalising the errors instead.
Established It depends, more than the literature admits, on coordination that is not technological. Third-party data reporting is the mechanism doing the work in every measurable success in this brief; it is a reporting obligation on employers and banks, not a system. Defaults are a statutory choice. Clearinghouses are deadlines and a queue. Each of the demonstrated wins is a piece of institutional plumbing that a government could install without buying anything.
Frontier What depends on it is more contingent than usually assumed. Future Civil Services depends on this brief’s evidence in the strict sense that its workforce projections assume a substitution effect this brief cannot find. Future Public Administration depends on it for whether delivery improves. Digital Citizenship depends on it for whether digital public infrastructure delivers what its business cases promise. And any fiscal plan that books savings from administrative automation depends on a conversion rate from time saved to output that nobody has estimated.
6 · Required experiments
Established The experiments that would settle the central questions are cheap, and their absence is a choice rather than a difficulty. Five, in ascending cost, with the natural experiments already running noted where they exist.
Established Measure output, not minutes. Instrument a caseload rather than a diary: cases cleared per week, decisions reversed on appeal, backlog age, error rate at quality-control sample. This is the study that does not exist, and constructing it requires no new method — every one of those numbers is already produced by the administrations in question for other purposes. Publish the override rate for every deployed system claiming human review, broken down by the groups the system over-refers; the United Kingdom’s Department for Work and Pensions has effectively done this and the result was informative, with caseworkers reversing referrals at between 14% and 77% depending on age band. Publish a denominator — an inventory independently verified rather than self-reported, which its own auditor has already established the current version is not. Frontier Audit the human baseline on the same measures used to audit the model; Amsterdam is the only case located where this was done, and it found the caseworkers biased too. Frontier Randomise. Assign comparable caseloads to algorithmic and manual triage and measure accuracy, speed and distribution of error. This is standard in medicine and essentially absent here.
Established One evaluation has come closest and it is instructive how close that is. The Business and Trade pilot randomly selected about 30% of its participants, stratified by directorate and grade, and blind-scored recorded tasks. That is not a randomised controlled trial — the majority were volunteers — and it still produced the only negative finding in the field. The cost of getting from there to a proper trial is one department’s willingness to withhold licences from a randomly chosen half of a directorate for three months.
Frontier Elicit forecasts before running the evaluation, because the pattern of forecasting error is itself a finding. The eleven-year follow-up of a randomised institution-building programme in Sierra Leone asked 126 experts to predict its results first. On physical infrastructure the mean forecast of 0.218 SD almost exactly matched the realised 0.204. On institutional change experts predicted 0.095 against a realised 0.062 SD, not significant after adjustment for multiple comparisons — and the local policymakers closest to the programme predicted around 0.25, roughly four times the truth. Expert judgement was well calibrated about the thing you can see and badly calibrated about the institutional effect, and proximity made it worse. Every claim in this field about what AI will do to public administration is an institutional forecast made by proximate parties.
Frontier The natural experiments already running have dates attached and are worth marking now. The United Kingdom has committed to a national digital exchange targeting £1.2 billion a year in savings, to one in ten civil servants in digital or technology roles by 2030 against a March 2025 actual of 5.2%, and to a consultation-analysis tool with a median review time of 23 seconds per response. These are the clearest testable commitments any government has made in this area. All three come from the implementing body and none is accompanied by a published measurement plan, which is the thing to watch: whether anyone marks them in 2030.
Established Negative results worth recording as experiments in their own right. New York City’s municipal business chatbot, launched in autumn 2023, told businesses that landlords could refuse housing vouchers, that employers could take a cut of workers’ tips, and that shops could refuse cash — the last in direct contradiction of a 2020 city law. Exposed by investigative reporting in March 2024, it stayed live for nearly two more years behind a disclaimer telling users not to treat its answers as advice; it cost nearly $600,000 to build and about half a million dollars a year to maintain, and was shut down in February 2026 as “functionally unusable.” Amsterdam’s welfare-fraud pilot performed every accountability step available and was stopped by the pilot outcome, not by the apparatus. And an independent evaluation of Chicago’s predictive-policing pilot found that 426 individuals on its Strategic Subject List were no more or less likely to become victims of homicide or shooting than 17,754 matched comparisons, with the apparent city-level decline part of a pre-existing trend. Three deployments evaluated honestly, three nulls or worse.
7 · Engineering requirements
Established The systems in this brief are not, technically, remarkable, and that is the most important engineering fact about them. The models involved are logistic regression, gradient boosting, rules engines, and in the most notorious case simple division. The general-purpose assistants now being deployed across administrations are commercial office products with a licence attached. Nothing on the critical path is a research artefact, which means that every difference in outcome between one deployment and another is a difference in process design, data supply and review capacity rather than in model quality.
Established The directly observed task decomposition is the single most useful object in the literature, because it is where self-report and measurement come apart. All figures below are from the Department for Business and Trade evaluation, October to December 2024, eleven tasks recorded on video and blind-scored one to five, three participants per condition:
| Task | Self-reported saving | Directly observed |
|---|---|---|
| Data analysis in a spreadsheet | +0.6 hours | Slower — 25:01 against 20:33; accuracy 1.5 against 2.7 |
| Slide production | ≈0 hours | 7+ minutes faster; quality 1.5 against 5.0 |
| Email writing | +0.2 hours | Marginally faster — 7:30 against 7:43; accuracy 4.3 against 4.0 |
| Summarising a report | +0.8 hours | Much faster — 12:37 against 41:34; accuracy 4.0 against 2.5 |
| Scheduling | — | −0.6 hours (net loss) |
Established Read the shape rather than the rows. The tool compresses a long document extremely well and with better accuracy than the human baseline; it degrades numerical work; and the self-reports do not track either direction. A department that measures adoption by asking people how much time they saved will get a number that is uncorrelated with what the tool did to its work. The same evaluation records that 22% of diarists reported hallucinations and that between 23% and 64% of users, depending on task, did not quality-assure the output or could not compare it against an alternative.
Established The engineering requirements that actually govern safety are unglamorous and mostly not about the model. Logging sufficient to reconstruct why an individual decision was made, years later, in a tribunal. Notice that reaches a person who does not check an online portal. A review path a caseworker can exercise, with the time budget to exercise it. Retention rules that survive audit. And — the requirement nobody builds — instrumentation that records the override rate and the outcome of every override, so the system can be evaluated on the thing that determines whether it is safe.
Established The technical choices that look most protective are weaker than they appear. Amsterdam’s welfare-fraud model used an explainable model class, excluded gender, nationality, age and postal-code proxies by design, and was trained on 3,400 prior investigations. Pre-deployment testing found it twice as likely to wrongly flag non-Western as Western nationals anyway; the training data was reweighted; and in a live pilot of nearly 1,600 applications the bias re-emerged in the opposite direction, the model flagged more applicants than intended, and it was no better than caseworkers at finding fraud. Excluding protected attributes did not produce a fair system, and the pilot — not the design review — is what discovered that.
Established Finally, the deployment inventory is an engineering artefact rather than a measurement, and should be read as one. The United States federal generative-AI count rose from 32 in 2023 to 282 in 2024 inside a total AI inventory rising from 571 to 1,110, then to 3,611 use cases across 56 agencies in the 2025 inventory published in April 2026, of which 445 are designated high-impact and 102 are a single commercial assistant. But in the 2024 snapshot 30% of generative-AI use cases were merely “initiated” and 27% were in acquisition or development — a majority not in operation — and 61% were mission-enabling back-office functions rather than government services (15%) or health and medicine (9%). Most of what is counted as artificial intelligence in government is not deciding anything about anyone.
8 · Adjacent technologies
Within this map the two closest neighbours are the two that share evidence with it. Future Civil Services holds the same three departmental evaluations and reads them as a labour question — whether machines substitute for administrative staff. This brief reads them as an output question. The split is clean because the studies answer neither: they measured minutes, so the workforce brief cannot get headcount from them and this one cannot get throughput. Digital Constitutional Systems holds the judicial record and owns what it means constitutionally; this brief takes from the same judgments only what they establish about how a system operated in practice.
Institutional Design is adjacent in the strong sense that it supplies this brief’s prior. The finding that a coordinated clearinghouse captured roughly 80% of the available welfare while every further algorithmic refinement was worth 0.6% to 2.7% is the sharpest available statement of where the gains in public decision-making are, and it was measured in a school-assignment system rather than an AI one. Collective Intelligence uses mechanism design as a tool where this brief meets it as a constraint.
Digital Citizenship supplies the randomised evidence on whether digital public infrastructure delivers, and owns the exclusion harms that follow when it does not; this brief takes the efficiency half of the same trial. Smart Cities found the same pattern in municipal instrumentation — success where a system gave an existing authority a better view of something it already did, failure where it was expected to substitute for a service that was not there. Megaproject Governance shares the accountability-gap problem on a longer timescale, and Civic Technology approaches the same institutions from outside them.
Adjacent outside this map: administrative law and the law of evidence; the economics of tax compliance, which supplies the strongest causal evidence here; market design, which supplies the strongest counter-evidence; and the audit profession, which has produced more usable findings on this subject than the computer-science literature has.
9 · Institutional requirements
Established The defining institutional fact is that nobody in this system has the job of finding out whether it works. Procurement buys, departments deploy, regulators check legality, auditors check process, and transparency registers record existence. No body anywhere in the chain is required to produce an output measure, and none does. That is why the entire evidence base consists of three departmental self-evaluations and a handful of pilot reports: the evaluation function is discretionary, and where it has been exercised, it has been exercised by the deploying department on itself.
Established The accountability apparatus exists, is growing, and has not been observed to stop anything. Transparency registers are post-hoc disclosure instruments: a system is entered once it is built, usually once it is live. Helsinki’s register is the clearest case — nine entries, six of them chatbots and two library systems, last modified in August 2024. Nothing in it decides anything about anyone. The critique made of these registers when they held three and five entries survives their growth: they are scoped to uncontentious municipal uses and omit the sectors actually implicated in algorithmic discrimination.
Established The United Kingdom’s mandatory register is missing its most consequential welfare system. The Universal Credit advances model — deployed since 2023, screening 1.4 million advances worth £0.8 billion in a single year — does not appear on the transparency hub, and the department has refused to confirm how many further counter-fraud models are in development.
Established Impact assessments are in the same position, and the responsible body’s own review is the evidence. Canada’s Directive on Automated Decision-Making is the longest-running binding instrument of its kind and has produced 39 published algorithmic impact assessments in seven years, heavily concentrated in two departments — thirteen from the immigration department and nine from the employment and social development department. Its fourth review, run in 2024–25, proposes extending the Directive to Agents of Parliament, who are currently excluded, and presents no findings on compliance and no evidence that an assessment has influenced a design or deployment decision; instead it proposes creating the compliance machinery, including departmental compliance reports signed by an assistant deputy minister. After five years, the body that owns the directive was proposing to start measuring whether anyone follows it. Frontier The one plausible case of an assessment producing a hard design constraint could not be verified at source.
Frontier The European backstop is real, is not yet in force, and moved further away this year. The AI Act classes public-benefit eligibility, law enforcement, migration and justice systems as high-risk and requires public-body deployers to complete a fundamental-rights impact assessment and notify a market surveillance authority — the only such regime in the world. Regulation (EU) 2026/1744 entered into force on 27 July 2026 and moved those obligations from 2 August 2026 to 2 December 2027, with systems already in use by public authorities pushed to 2 August 2030, and permits the rights assessment to be discharged by cross-reference to a data-protection assessment. As this brief is written, the Act imposes no high-risk obligations on any public administration system. The prohibitions — social scoring, and predicting individual criminality solely from profiling — are in force, and that “solely” is doing a great deal of work.
Frontier What is actually maturing is individual and judicial rather than collective and administrative. The Court of Justice has held that a score can itself be the automated decision, and that a controller must describe the procedure and principles actually applied well enough that a person can understand which of their data were used and how varying it would have changed the outcome, while ruling that trade secrecy is not a ground for refusal. A counterfactual explanation owed to an individual on demand is a more powerful instrument than a register entry owed to the public in general, because the person harmed can enforce it. The constitutional development of that doctrine belongs to Digital Constitutional Systems; what belongs here is that it is the only instrument in the field with a demonstrated enforcement mechanism. Frontier The United Kingdom has moved the other way, replacing the general prohibition on solely-automated decisions with a risk-based approach that substantially widens what is permitted outside special-category data.
Frontier The institution that does not exist is an evaluator with the standing to withhold. There is no body that can decline to authorise a deployment for want of an output measure, no equivalent of the trial requirement that precedes a licensed medicine, and no procurement rule anywhere located that makes payment contingent on a measured administrative outcome. The consequence is visible in the adoption record: a technology is bought, counted, registered, and never marked. That is the shape of isomorphic mimicry — the documented pattern in which institutions adopt reforms that enhance external legitimacy “even when they do not demonstrably improve performance,” producing a capability trap in which reforms are constantly adopted and capability never improves. The diagnosis was written about public financial management reform in developing countries. It fits the AI adoption record in wealthy administrations exactly, and the authors of that diagnosis are explicit that their proposed remedy has no outcome evidence behind it either.
Frontier Procurement is the only chokepoint in this chain that acts before the system exists, which is why the clauses are multiplying and why they still miss the question. Model contractual terms for public buyers now circulate in several jurisdictions — a voluntary European template, the American acquisition memorandum as a binding instruction to contracting officers — and their content is consistent: data rights, limits on training a supplier’s models on the state’s records, documentation duties flowed down to subcontractors, monitoring, exit and portability against lock-in. These are good clauses, and they regulate the supplier relationship. None located makes acceptance, payment or renewal contingent on a measured administrative outcome. Handwave The likeliest route by which assurance acquires teeth is the dullest one — a contracting officer withholding a milestone payment for want of a monitoring report — and no instance of that has been located, in either direction.
10 · Ethical & societal considerations
The first ethical question in this subject is evidential, because almost everything favourable that is known about AI in government was reported by a party with an interest in the answer. Established The three most-quoted figures in the field — 26 minutes, 95 minutes, and a 3–13% GDP opportunity — come respectively from a government body evaluating its own pilot, a government body evaluating its own pilot, and a consultancy that sells the implementation. The most-quoted national savings figure comes from a government press office citing a government-aligned advocacy foundation. The most-quoted evaluation of a flagship government app returns 100% on two separate favourable items. This brief applies to those the same rule the corpus applies to a fusion company counting its own funding, and marks every one of them in the reading list. The exception, and the reason the field is not evidence-free, is the small set of sources reporting against interest: a department publishing a null result about its own pilot, an audit office demolishing its own government’s savings claim, and a supreme administrative court condemning its own case law.
Established The distributional finding is the most consistent thing in the evidence base. Risk selection over-selects the poor, migrants and the already-marginal, in jurisdictions sharing no legal order: a tax administration using nationality as an automated risk indicator; a student-finance body whose neutral-looking criteria disproportionately selected students with migration backgrounds; a welfare department referring non-UK nationals at 2.27 times the baseline rate. Frontier But one audit that traced selection through to outcome found the disparity did not propagate: the higher inspection chance did not translate into disproportionately higher benefit terminations. The harm was real and is being compensated at about €80 million, and it was a procedural harm rather than a differential-outcome harm — a distinction almost never preserved in secondary accounts, which changes what the remedy should be. The audit was commissioned by the body it audited.
Established The sharpest ethical finding here concerns the fallback rather than the machine. When Amsterdam cancelled its welfare-fraud model for bias, screening reverted to human caseworkers — whom the city’s own analysis had already measured as biased in a different direction, against Dutch nationals and women. The apparatus audits the algorithm and leaves the humans unaudited, so cancelling a model can substitute one uninspected bias pattern for another and be recorded as a win. There is no plan to audit the caseworkers.
Frontier The opportunity-cost question is not usually posed as one, and the mechanism-design evidence poses it sharply. If coordination captures roughly 80% of the available welfare in the one case anyone has decomposed, and algorithmic refinement captures a few per cent, then money spent on models in an administration that has not yet fixed its queues, deadlines and data-sharing is being spent on the wrong margin. There is a further, unwelcome result in the same literature: in a hand-coded study of 4,700 public projects across 63 federal organisations, a bureaucratic autonomy index was robustly positively correlated with project initiation, completion and completion rate while an incentives-and-monitoring index was robustly negatively correlated with all three. The published summary reports directions rather than coefficients and this brief attaches no magnitude to either. If the sign generalises, a technology whose principal administrative use is monitoring is being deployed against the grain of the only large-scale evidence on what makes a bureaucracy complete things.
Established And the deepest issue is not technical at all. Reversing the burden of proof — making a person disprove an accusation the state generated automatically — is a legal design choice. Automation makes it cheap at scale; it does not make it necessary.
What this brief cannot resolve, stated as an obligation rather than an apology. Whether the three evaluations of the same product differed by design alone or also by version, licence tier or training could not be established. Whether the conversion rate from time saved to administrative output is positive, zero or negative is unknown, because it has never been measured. Whether the New York decomposition generalises beyond school assignment is unknown; it has been performed once. Whether India’s ration-card cancellations number 2.06 crore, as the government states while denying the identity system was the cause, or 3 crore, as counsel asserted in the Supreme Court, is unaudited on either side, and both are given here without a preference. No independent outcome evaluation was located for any deployment of the main exported identity platform, in any of its fourteen production countries. And nobody has published a defensible estimate of whether machine-readable rules improved decision accuracy or changed outcomes for any claimant population — a gap that belongs to Digital Constitutional Systems and is recorded here because it bears on the same question.
11 · Civilizational implications
Established The civilisational change is that scale has removed the friction that used to limit how many wrong decisions a state could make. 866,857 compliance interventions in one Australian programme; 22,427 automated fraud determinations in under two years in one American state, 93% of them overturned; 1.5 million people profiled in Poland; over 13 million scores computed monthly in one French system; 270,000 people on one Dutch blacklist; 1.4 million advances assessed in a single British year. The error rate does not have to rise for the error count to become a national scandal. An administration that could previously make ten thousand mistakes a year can now make a million, and the appeals system behind it was sized for the former.
Established The second-order effect is the one the Australian remediation arithmetic exposes: at sufficient scale, unwinding an error becomes more expensive than the state can bear, and the institutional response is to legalise it. That is a new thing in public administration and it has not been reckoned with. It suggests the decisive governance question is not whether a model is accurate or explainable but whether the volume of decisions it generates is within the capacity of the review system behind it — a constraint nobody currently computes before deployment.
Established The general principle this case illustrates is about where welfare comes from, and it is not confined to government. In the one decomposition anybody has performed, coordination captured about 80% of the available gain and algorithmic sophistication captured between 0.6% and 2.7%. The corresponding result in auctions is that market structure, not mechanism, set a fourfold spread in outcomes. If that pattern is general, then a civilisation’s returns to computation in its collective decisions are bounded by how well it has coordinated first, and a great deal of what is spent on making public decisions cleverer is being spent on a residual that has already been measured and found small. That is a claim about the shape of the opportunity, not a claim that the residual is zero, and it rests on two decompositions in two sectors.
Frontier There is also a legibility hazard that this corpus has documented elsewhere and that applies directly here. The most influential governance index ever built caused over 70 countries to form regulatory reform committees oriented specifically to its indicators, was gamed by consultancies paid to move rankings, and produced fifteen-year swings of more than forty places for 111 countries and more than seventy-five places for 35 — volatility no plausible account of underlying institutional change explains. When you make an institutional target legible and scored, you get the score rather than the institution. Any regime that measures administrative quality by a published algorithmic metric should expect the same, which is an argument for measuring outputs an agency cannot easily manufacture — cases cleared, appeals upheld — rather than compliance artefacts it can.
Frontier Against all of that, the genuinely transformative results in this area are undramatic and already proven: third-party reporting, pre-population and sensible defaults reduce both error and burden without predicting anything about anyone. Handwave Visions of comprehensively algorithmic government rest on an evidence base that, on the record assembled here, does not exist in either direction.
12 · Timelines
Established What already happened, because this timeline usually starts too late. The instruments predate the technology. Canada’s Directive on Automated Decision-Making has been binding since 2019 and has produced 39 published algorithmic impact assessments. The Dutch risk-indication system was struck down in February 2020, the Australian royal commission reported in July 2023, and the United Kingdom’s transparency standard became mandatory across central government in 2024. The three departmental evaluations that constitute the entire evidence base for general-purpose AI in a civil service ran between October 2024 and March 2025, which is to say that the evidence is younger than most of the policy built on it.
Frontier 2026 to 2027: the measurement window, if anyone opens it. Deployment counts are rising faster than evaluation capacity — a 105% increase in reported United States federal use cases in one year against an inventory its own auditor found unreliable. Whether any administration publishes an output measure in this period, rather than another time-saved survey, is the single most informative thing that could happen to this field, and nothing currently scheduled requires it.
Frontier December 2027 and August 2030: the European high-risk regime, twice deferred. Regulation (EU) 2026/1744 entered into force on 27 July 2026 and moved Annex III high-risk obligations from 2 August 2026 to 2 December 2027, with systems already in use by public authorities pushed to 2 August 2030. As this brief is written the Act imposes no high-risk obligations on any public administration system anywhere. Handwave Any date given for when it will bite should be discounted by the fact that this deadline has already moved once, in the direction of later, by sixteen months.
Frontier 2030: the United Kingdom’s stated targets fall due. One in ten civil servants in digital or technology roles, against 5.2% in March 2025; a national digital exchange saving £1.2 billion a year. Both are the implementing body’s own targets. Neither has a published measurement plan, and the corpus’s experience of decadal reform targets is that the marking exercise is the part that does not happen.
Speculative The 2030s: either the evidence base acquires an output measure or it does not. If override rates, denominators and case-throughput data become ordinary published statistics, this field gets an evidence base and the questions in this brief become answerable. If they do not, the literature continues to consist of scandals plus self-assessment, and the deployment decisions continue to be made on vendor and departmental self-report. Nothing technical determines which of those happens.
Handwave Beyond that, no useful forecast. The durable questions — who bears the burden of proof, and who pays when the state is wrong at scale — are older than computing and will outlast any particular technology. Projections of end-to-end automated administration assume a remediation capacity nobody has demonstrated at any scale.
Speculative The fork, and it is decidable rather than mysterious. Either override rates, denominators and outcome measures become ordinary published statistics — in which case this field acquires the evidence base it currently lacks — or they do not, and the literature continues to consist of scandals plus self-assessment. Nothing technical stands in the way: the agencies running these systems already know how many decisions were made and how many were overturned. Frontier On present trends the second is the better bet, because the three best-designed evaluations available all measured minutes rather than output, and the one that looked hardest concluded it could not identify how the saved time was spent.
- 2 yr: Frontier The first hard compliance date in public-sector AI assurance has now fallen due — high-impact federal systems brought into line with the minimum practices or withdrawn — and three quantities decide what it meant: how many systems were declared high-impact, how many practices were waived, and how many were actually switched off. Handwave None of the three has been located in published form, and the corpus’s experience of deadlines attached to self-declared inventories is that the inventory adjusts to the deadline rather than the estate adjusting to the rule.
- 10 yr: Frontier Either an output measure, an override rate and an incident duty become ordinary published statistics, or assurance settles at its present equilibrium of assessments filed, audits unread and rates uncomputable. Speculative The European high-risk obligations bite on public administration or are deferred a third time.
- 25 yr: Speculative The durable questions — who bears the burden of proof, and who pays when the state is wrong at scale — are settled by liability rules rather than by any assurance framework, because that is how every previous administrative technology was eventually disciplined.
13 · Technology tree & dependencies
- Depends on Nothing on this map. Nothing in this topic waits on a result produced by another brief here: the models are ordinary, the products are commercial, and nothing on the critical path is a discovery. The constraints are measurement and administrative capacity, and they are recorded on the row below rather than described and dropped.
- Requires (not on this map) An output instrument — cases cleared, error rate, backlog age — because the entire evaluation literature measures self-reported time on task and nothing else, and no deployment can be judged without one. Published override rates, without which “a human reviews it” is unfalsifiable. A verified inventory of deployed systems, without which no rate of anything can be computed, and which the responsible auditor has established the current self-reported version is not. A review path with the capacity to reverse decisions at the rate the system generates them — the binding constraint, priced by the Australian remediation arithmetic at roughly five hundred times the cost of legalising the errors instead. Third-party data reporting, which is the mechanism doing the work in every measurable success here and is a statutory obligation rather than a system. And a funded liability for being wrong at scale. And a compulsory incident-reporting channel, because every other safety-critical public function combines a duty to report occurrences with a denominator of operations, and this one has neither — the only repositories are voluntary and built from media coverage, so the observed failure record is the distribution of investigative attention rather than of failure. None is a research result; five are institutional, one is fiscal, and one is a reporting rule.
- Enables Accurate, fast, low-burden administration at scale — conditionally, and only where the constraints on the row above are met. No typed enabling edge is claimed, because on the evidence assembled here the enabling relation is unproven in both directions: no published study of a government AI deployment measures whether output improved.
- Adjacent AI Governance as the opposite subject; Digital Constitutional Systems for the constitutional and rights analysis of automated decisions; Future Civil Services for the workforce half of the same evaluations; Institutional Design for the coordination-versus-optimisation decomposition this brief leans on; Digital Citizenship for the randomised evidence on digital public infrastructure; and Smart Cities for the same pattern in municipal form.
14 · Common misconceptions & speculative claims
“AI saves civil servants about 26 minutes a day.” Established Established as a self-report; the causal claim is unsupported. The figure is a survey answer from a 20,000-person trial with no comparison group, in a report that says outright it could not identify how the saved time was spent. The same product measured against a comparison group at another department gave 19 minutes. The same product evaluated with directly observed, blind-scored tasks produced no evidence of a productivity improvement at all, with users measurably slower and less accurate at spreadsheet analysis. Quote the range and the design, or quote nothing. This brief does not resolve one residual conflict and states it: the three evaluations ran in overlapping periods and it could not be established from the material consulted whether they used the same product version, licence tier or training provision, so the design explanation is the best available one rather than the only possible one.
“Pennsylvania proved generative AI saves 95 minutes a day.” Handwave One hundred and thirty-six exit-survey respondents from 175 volunteers across fourteen agencies, self-reporting, in a sample the report explicitly calls unrepresentative, with an acknowledged undercount of people who stopped using the tool and stopped answering. The report is candid and useful. The statistic is not.
“Agencies are deploying thousands of AI systems.” Established for the count, Frontier for the interpretation. The 2025 United States federal inventory reports 3,611 use cases, and it is a reporting artefact. In the prior year’s examination a majority of generative-AI use cases were merely initiated or in acquisition, 61% were back-office, only five of twenty agency inventories were comprehensive, duplicates and non-AI systems were included, one agency reported 375 uses privately against 33 publicly, and the Pentagon and Intelligence Community do not report at all.
“The clever algorithm is what delivers the gain.” Established In the only case where anyone has separated the two, it delivered between 0.6% and 2.7% of it. Coordination — a single deadline-respecting clearinghouse replacing an uncoordinated scramble — delivered about 80%. The same conclusion arrives from spectrum auctions, where a fourfold spread in per-capita revenue across four European countries in one year tracked how many serious bidders turned up rather than which mechanism was used. This is the correction that matters most for anyone proposing a model to improve a public decision, because it says the thing being optimised is the small residual.
“Estonia has a robot judge deciding small claims.” Established It does not and never did. The Ministry of Justice stated on the record on 16 February 2022 that it “does not develop AI robot judge for small claims procedure nor general court procedures to replace the human judge,” and called the March 2019 reporting misleading. What Estonia actually automates is the order-for-payment procedure — about half of civil cases — plus transcription and anonymisation. It remains one of the most-cited examples in this field and it did not happen. This brief separately notes that no evaluation of Estonia’s or Singapore’s government AI programmes could be verified from the material consulted; they may be named as the standard exhibits and no figure from either appears here.
“Digital government has saved country X billions.” Established as a claim, Handwave as a figure, and there is a clean natural experiment on how such figures survive audit. India’s headline claim of 3.48 lakh crore rupees saved by identity-linked benefit transfer between 2009 and 2024 is published by the government’s own press office, which names its source as a quantitative assessment by a private, government-aligned advocacy foundation whose stated purpose is promoting a national development programme and whose output is policy papers and ministerial commentary. It is not an auditor, a university or a statistical agency, and the methodology is not reproduced. When an audit office did test an earlier claim of the same shape — 21,552 crore rupees from cooking-gas subsidy transfer — it found about 92% of it attributable to the collapse in crude prices and 7.6% to the identity system.
“Digital identity could unlock 3–13% of GDP.” Handwave It is a consultancy’s general-equilibrium model output for seven countries, published in 2019, assuming roughly 70% adoption by 2030 plus accompanying digital infrastructure. The consultancy says in its own text that these are not forecasts, that “realizing this value is by no means certain or automatic,” and that not all the potential value may translate into GDP. There is no independent replication of the range and no ex-post study measuring whether any deploying country landed inside it. The figure is a model output from a firm that sells digital-transformation work, and it circulates as though it were an observation.
“The evaluation found the programme was a success.” Established Check who wrote it. The principal evaluation of one widely praised national government app was produced by a body whose institutional purpose is promoting digital government, published on a platform funded to promote digitalisation, and reports that 100% of respondents had never experienced corruption through the app and 100% rated its social impact positive or very positive, alongside 96.8% agreeing it simplified identification, and identifies no major shortcomings. A survey instrument returning 100% on two separate items is measuring its own respondent selection. The same standard applies domestically: a department that publishes figures showing claimants aged 66 and over referred for investigation 49 times more often than the baseline age band and less than a quarter as likely to be correctly rejected, and then concludes from its own data that there are “minimal concerns of discrimination”, is an interested party assessing its own product.
“Robodebt was artificial intelligence.” Established It was data-matching and division — no model, no training data, no score. Opening a discussion of algorithmic government with it, as almost everyone does, miscasts the field’s own lead example and relocates a legal-design failure into a technology.
“SyRI was struck down for discrimination, under the GDPR.” Established It was held incompatible with Article 8(2) of the European Convention on the fair-balance limb. The court expressly declined to decide the automated-decision question, held that risk profiling is not per se contrary to Article 8, accepted a pressing social need, refused the order to disclose the risk models, refused the order to destroy the data, issued a declaration running only with respect to the successful claimant organisations, and recorded that the system had been used in just five projects. It is also not evidence that algorithmic fraud detection does not work: the court condemned opacity and disproportionality, not inaccuracy, and the government’s refusal to disclose is precisely why the effectiveness question is unanswered rather than answered negatively.
“Michigan’s 93% error rate is an Auditor General finding.” Established It is the agency’s own review of 22,427 determinations, later recorded as such in the state Supreme Court’s opinion; the Auditor General’s 2016 performance audit expressly excluded fraud identification from its scope. The misattribution is widespread, including in advocacy sources. Established “Poland’s profiling system was struck down for being discriminatory.” The Constitutional Tribunal’s objection was formal: the scope of data used should have been fixed by Parliament rather than by ministerial regulation.
“Human review is rubber-stamping.” Frontier Not an established empirical finding. The best-powered public-sector study found civil servants less likely to follow algorithmic than identical human advice, and one department’s caseworkers override its model for the worst-served age band 77% of the time. The defensible statement is that oversight is active and unaudited. Established And the direction of judicial travel is not one-way either — in September 2025 an Austrian court annulled a data-protection authority’s ban on a public employment-service profiling algorithm, holding adviser involvement “not merely symbolic, but essential”.
“Transparency means publishing the code.” Established France’s family-benefits agency released its source code in January 2026 after four years of litigation, and the Court of Justice has held that “the mere communication of an algorithm” is not a sufficient explanation. What a person is owed is an account of how their own data produced their own result. Handwave And two cases should stop being cited as they are: the Danish report says in terms that it could not determine impact and was refused the data needed for bias testing — a finding about opacity routinely re-reported as a finding about discrimination — while nothing in the Indian audit supports a quantified welfare-exclusion claim. What that audit does establish, and what is worth citing, is that the operator of a 1.4-billion-person identity system had no system for analysing the causes of its own authentication failures.