1 · Concept overview

A civil service is a strange institution to hold an opinion about, because almost everything asserted of it is asserted without measurement. It is said to be too big, too old, too generalist, too slow, insufficiently digital, insufficiently commercial, captured by its own procedures, and about to be replaced by software. Reform plans arrive on roughly a decadal cycle carrying those diagnoses and a numbered annex of commitments, which makes the plans unusually easy to mark. Marking them is what the first version of this brief did, and the marking held: on the two things that can be counted cleanly in the United Kingdom record — headcount and senior pay — the reform did not merely fail to arrive, it reversed.

This version keeps that spine and adds the four questions the marking exercise cannot reach. First, what happens when a permanent bureaucracy meets general-purpose automation: not what vendors and ministers say will happen, but what the three published evaluations of AI assistants inside civil services actually measured, which is less than is usually reported and in one case is nothing at all. Second, what became of the digital service units built to fix these institutions from inside — 18F, the United States Digital Service, Canada's CDS, the UK's GDS — a record that turns out to be about org charts and executive orders rather than about performance. Third, the demographic structure, which is not a forecast about the 2030s but a fact about the present: 7.9% of the United States federal workforce was under 30 in January 2026 against 19.4% of the national labour force, and in federal information technology there are roughly twenty times as many staff over 50 as under 30. Fourth, the politicisation question, which most writing treats as a values argument and which has a quantitative literature pointing consistently one way.

The scope boundary with the sibling slot is the unit of analysis. This brief owns the workforce and its rules: headcount, pay structure, appointment and tenure protections, age and skill composition, contracting-out of the labour, and the evidence on whether machines substitute for administrative staff. Future Public Administration owns programmes and delivery: what was built, whether it worked, and what the project record shows. Where they meet — the Phoenix pay system is the clearest case, since it is simultaneously a project failure and a labour-relations catastrophe — this brief takes the workforce half and the sibling takes the delivery half.

The brief is deliberately not Canada-specific and deliberately not United Kingdom-specific. It draws on both because both have supreme audit institutions that publish, which is the same reason it draws on Australia's Robodebt Royal Commission and on United States accountability offices. That selection is itself a bias worth naming at the outset: the countries with the best-documented civil service failures are the countries with the best-functioning auditors, and the absence of a comparable record elsewhere is not evidence of a better record.

2 · Current scientific position

Frontier The best-designed evaluation of an AI assistant in a civil service found no productivity gain, and the two evaluations that found large gains had no comparison group. This is the single finding that most changes how this subject should be read, and it is available because three United Kingdom departments evaluated the same product with three different designs in overlapping periods. The Department for Business and Trade ran Microsoft 365 Copilot across 1,000 licences from 1 October to 30 December 2024, about 70% volunteers and 30% randomly selected stratified by directorate and grade, and included eleven directly observed tasks recorded on video and blind-scored for quality and accuracy. Its conclusion is stated without hedging: “The evaluation did not find evidence that time savings have led to improved productivity, and control group participants had not observed productivity improvements from colleagues taking part in the M365 Copilot pilot.”

Established The famous 26 minutes a day is a survey answer. The cross-government experiment behind it covered 20,000 users across twelve organisations, with survey responses from 7,115 and telemetry from 14,500, and reported 26 minutes saved per working day, annualised to 13 days a year, with 17% reporting no clear saving. It had no comparison group, its own limitations section warns that the figures are self-reported and should be weighed against literature on what such metrics are worth, and on the question that matters it is explicit: it was not possible to identify how the saved time was spent. When the Department for Work and Pensions re-ran the same product across 3,549 licences with a genuine comparison group — 1,716 users against 2,535 non-users, adjusted for demographics, role and prior AI experience — the figure fell to 19 minutes a day. Adding a comparison group cost roughly a quarter of the effect. Removing self-report entirely, at DBT, cost the rest.

Established The observed-task data are more useful than any headline and they say the tool is good at one clerical job and bad at another. On blind-scored tasks, summarising a report took 12 minutes 37 seconds with the assistant against 41 minutes 34 seconds without, and accuracy rose, 4.0 against 2.5 on a five-point scale. Spreadsheet data analysis took 25 minutes 1 second with against 20 minutes 33 seconds without, and accuracy fell, 1.5 against 2.7. Scheduling was a net loss of about 0.6 hours. Self-reported savings tracked none of this. The observed-task sample was three per condition and the report says so, so these are indicative rather than estimated effects — but they are the only direct measurements in the literature, and they predict something specific: substitution in précis and briefing work, and no substitution whatever in analysis.

Handwave The 95 minutes a day that circulates in North American coverage. Pennsylvania equipped 175 employees across 14 agencies with ChatGPT Enterprise for twelve months to March 2025; 136 completed the exit survey and reported an average 35 minutes a day using the tool and 95 minutes a day saved. The report is honest about what that is: participants “were not a representative sample of Commonwealth employees,” the share of non-users is “likely an underestimate” because people who stopped using the tool also stopped answering surveys, and a pre-pilot expectation of six use cases per employee came out at three. Its own conclusion — “ChatGPT is not a substitute for the nuance and experience employees bring” — is the sentence that should travel, and it is the one that does not.

Established The demographic constraint is not a forecast and it is not fixable by a recruitment target. The United States federal civilian workforce stood at 2,035,344 in January 2026, the smallest in at least fifteen years, down 12% from 2,313,216 in September 2024 — the largest annual fall since the 1990s. Over that period 386,826 people separated, including 136,822 through deferred resignation and 10,436 through reductions in force, against 122,598 hires in a service that normally hires about 250,000 a year. And the share under 30 fell, from 8.9% to 7.9%, against 19.4% for the United States labour force, while the share aged 50 or over stood at 40% against 32.8%. A voluntary-exit programme combined with a hiring freeze makes a workforce older, not younger, because it widens the exit and closes the entrance. The composition effect is the finding; the headcount is the headline.

Established The digital service units did not fail; they were reorganised. 18F was founded on 19 March 2014 with fifteen staff, peaked above 250 in 2018, ran on cost recovery so that client agencies had to choose to pay for it, and was at roughly 90 distributed staff when all of them were terminated by email at 1 a.m. on 1 March 2025, having been locked out of their systems at midnight, as part of a reduction in force identifying the office as “non-critical”. It had built Login.gov, cloud.gov, the US Web Design System, analytics.usa.gov and COVIDtests.gov. The United States Digital Service was renamed in January 2025, lost about a third of its staff to an anonymous email on 14 February, and on 25 February twenty-one of its engineers, data scientists and designers resigned together rather than, in their words, “use our skills as technologists to compromise core government systems”. Canada's CDS was not abolished but relocated, moving in July 2023 from the Treasury Board Secretariat — a central agency — into Employment and Social Development Canada, several layers inside a service-delivery department of 32,000 staff. Not one of these decisions cites an evaluation.

Established The politicisation question has a quantitative literature and it points one way. A meta-analysis of 96 peer-reviewed studies across more than 150 countries finds merit-based recruitment, appointment, promotion and tenure protection positively and consistently associated with government performance and negatively associated with corruption. A one-standard-deviation increase in politically appointed managers is associated with a 51–63% increase in the probability of a non-competitive contract award. Political appointees average 18 to 24 months in post, which caps how much programme-specific knowledge they can acquire. And there is no evidence of partisan cycling among career civil servants against significant cycling among appointees — the finding that most directly contradicts the premise that a career service is a partisan actor requiring correction. This literature is correlational and this brief says so; it is not absent, and it does not point both ways.

Frontier Against it, the instrument exists and has been used at about a sixth of its announced scope. As of June 2026 roughly 8,000 career federal positions had been moved into Schedule Policy/Career by executive order, about 97% at GS-15 and above, concentrated in the budget office and covering division heads, chief information officers, senior legal and policy officials and regulation writers. The Office of Personnel Management's working estimate was 50,000; earlier projections reached 200,000. Those in the schedule lose Civil Service Reform Act protections, Merit Systems Protection Board appeal rights, the ability to challenge reclassification, and recruitment, retention and relocation incentives; agencies gain at-will removal. The gap between 8,000 and 50,000 cuts both ways: implementation is far short of the announcement, and the same instrument that moved 8,000 can move the remainder without further authority.

Established And the outsourcing question turns out to be a question about records rather than about value for money. The Auditor General of Canada's estimate that ArriveCAN cost $59.5 million is the auditor's reconstruction, not the government's figure, because 18% of tested contractor invoices did not document which project they belonged to: “We were unable to determine a precise cost for the ArriveCAN application because of poor documentation and weak controls.” External resources billed an average $1,090 per day against $675 for the equivalent Government of Canada position, a 61% premium, and a two-person intermediary firm took $19.1 million of the total. In the United Kingdom the National Audit Office reported in November 2025 that four sources give four figures for the same year's consultancy spend — £1.34bn, £1.36bn, £1.68bn and £2.23bn — and concluded that the government “does not have a clear picture of how much is being spent on consultants”. The state cannot currently say how much of itself it has contracted out.

3 · Frontier questions

The genuinely open questions in this subject are narrower than the discourse and mostly concern measurement rather than possibility. Separating them from the questions that merely sound open is most of the work.

Frontier The real frontier: does automation substitute for administrative labour, or add to it? Nothing located answers this. Every recent evaluation of AI assistants in a civil service measures time on task. None measures cases cleared, decisions issued, backlog reduced or staff required. The cross-government experiment states that it could not identify how saved time was spent; the DBT evaluation states that productivity “was not a key aim of the evaluation and therefore limited data was collected”; the DWP study measured minutes. There is, on the evidence consulted for this brief, no published study of an AI deployment in a government agency that measures output. The policy edifice built on top — a £45 billion opportunity, a target of one civil servant in ten in a digital or technology role by 2030, a £1.2 billion a year saving attributed to a single platform — rests on time-on-task self-reports, which is the same class of evidence this brief already refuses to accept from savings announcements.

Handwave The claim that computerisation of government historically just added a layer. This is the sceptic's version and this brief will not assert it, because the one empirical study located points the other way. Lee and Perry's test of the productivity paradox in United States state governments, published in 2002, found that “IT investments by state governments have positive and significant effects on the measure of productivity, gross state product,” with larger returns where a chief information officer structure existed. That is one study, one country, one tier of government, twenty-four years old, and only its abstract was accessible — the effect sizes, the number of states, the years and any employment effect could not be read and are recorded as a gap. The honest position is that the historical question is under-researched rather than settled, and that the sceptical intuition is currently an intuition.

Frontier What is genuinely uncertain about the demographic cliff is not whether it arrives but whether it is survivable through composition change. An institution with twenty times as many information-technology staff over 50 as under 30 has already lost the succession; the open question is whether a technical function can be sustained by a smaller number of people with better tools, by contracting, or not at all. The three routes have different failure modes and no country has run the experiment cleanly, because everywhere the demographic change, the automation and the contracting arrived together.

Frontier Whether merit protections can be removed and restored, or only removed. The comparative literature establishes that merit systems co-occur with performance and with lower corruption, and that reducing protections is predicted to drive the most motivated staff out first. What it does not establish is reversibility. The theoretical work implies a selection effect — if protections go, the people who would have most objected leave, and the remaining workforce is the one that accepted the terms — which would make restoration much slower than removal. No case of a modern civil service losing tenure protections and rebuilding them with the same people has been examined here, and the current American case is running in real time.

Frontier Whether the state can recover a capability it has contracted out. The National Audit Office found knowledge transfer from consultants to be systematically neglected rather than merely imperfect: officials typically wait until a project ends before discussing what was learned; consultants' intellectual property protections limit sharing; departmental staff turnover breaks continuity; and contracts specify learning obligations ambiguously. Its formulation is the diagnosis — “Knowledge transfer should not just be a transfer of documents at the end of a project but should be an active process throughout the engagement”. If that is right, contracting-out is not a reversible decision at the margin but a ratchet, and the empirical question of how far a department can climb back is open and, as far as this brief could establish, unstudied.

Established What is not open, stated plainly because it is repeatedly presented as open. Whether AI tools help with summarisation: they do, measurably, with better accuracy than the unaided baseline in the one direct test located. Whether they help with numerical analysis: they did not, in the same test, and made accuracy worse. Whether merit systems are associated with better government: the association is established across 96 studies and 150 countries and the argument is about mechanism, not existence. Whether digital service units were closed because they failed: no evaluation is cited in any of the closures examined. And whether governments know what they spend on consultants: they do not, and their own auditor says so.

Frontier One frontier that is more open than it looks: the reliability of the adoption statistics themselves. United States federal agencies reported 3,611 AI use cases across 56 agencies for 2025, up 105% from 1,757, with 445 designated high-impact. But in a prior review of 23 agencies, only 5 of 20 submitting inventories provided comprehensive information, with duplicate entries, references to non-AI technology and improperly included research; one agency reported 375 uses to the auditor against 33 on its published list. The Pentagon and the Intelligence Community are exempt from reporting altogether. Nobody currently knows how much AI is deployed inside the United States government, and the number that is quoted is a compliance artefact whose error is unmeasured in a known direction.

4 · Technological bottlenecks

Established The first bottleneck is the age structure, and it binds before anything else because it cannot be relieved quickly in either direction. On a Partnership for Public Service analysis of United States federal data, 6.2% of federal employees were under 30 against 23.9% nationally, 45.1% were 50 or over against 33.5%, 18.2% were retirement-eligible — and in information technology specifically there were twenty times as many staff over 50 as under 30. By January 2026 the under-30 share had fallen to 7.9% on a comparable measure. A workforce with that shape cannot recruit its way out on any relevant timescale, because the people who would train the entrants are the people leaving.

Established The second is that the standard cost-reduction instrument makes the first bottleneck worse. Deferred resignation programmes and hiring freezes remove people who were closest to leaving and close the only channel that lowers the average age. The United States ran both simultaneously through 2025: 136,822 deferred resignations and 10,436 reductions in force against 122,598 hires in a service that normally hires around 250,000, and the workforce got smaller and older at the same time. This is not a criticism of the policy's objective; it is an observation that headcount reduction and capability preservation are pursued with the same instrument and it can only do one of them.

Established The third is pay compression, which is the mechanism by which the age structure becomes a skills structure. In the United Kingdom, doctors fell from the 95th to about the 90th percentile of the national earnings distribution between 2007 and 2023, secondary teachers from the 87th to the 81st, and entry-level police officers from the 34th (2014) to about the 26th. The public sector paybill is about £270 billion a year and closing the public–private gap back to its 2019 position would cost roughly £17 billion above current plans. The gap is widest in London and the South East — which is to say, widest where the alternative employers are. And the compression reaches the deferred half of the package too: 13% of those earning £10,000–£16,000 opt out of the pension against 6% of those earning over £31,000.

Established The fourth is that compression shows up as fill rates rather than as quit rates, which is why it is chronically under-detected. Fewer than 50% of District Judge vacancies were filled in recent rounds; secondary teacher training recruitment ran well below target; prison service turnover was 13% in 2023; nurse and midwife vacancies stood at 8%; and 25% of the senior civil service changed role, changed department or left in 2022-23. An organisation watching attrition sees a manageable number. The vacancy that is never filled does not appear in an attrition statistic at all.

Frontier The fifth is contracting dependence, and the binding constraint is that it cannot be measured. Four sources give four figures for United Kingdom central government consultancy spend in 2022-23 — £1.34bn, £1.36bn, £1.68bn and £2.23bn — with an average variance of £270 million a year across six years, because departments use different definitions, bundle consultancy with professional services and contingent labour in the same contracts, and arm's-length bodies report to different standards. A dependency whose size is uncertain by a factor of 1.7 cannot be managed down to a target, and targets have nonetheless been set: halving spend in 2025, over £700 million a year of savings by 2028-29.

Established The sixth binds specifically on automation and is rarely named: quality assurance capacity. Between 23% and 64% of users of an AI assistant, depending on task, did not check its output or could not compare it. The constraint on deploying these tools at scale in an administration is not licences or model capability; it is the number of people with enough domain knowledge to know when an output is wrong, which is the same population the age structure and the pay structure are removing.

Established What is not a bottleneck, stated plainly. Model capability is not the binding constraint on AI in government; on the only direct measurements available the tool already outperforms the unaided baseline at the task it suits. Licence cost is not the constraint. Political appetite is not the constraint — the targets are set and the money is committed. The constraints are an ageing technical workforce, a pay structure that cannot compete for its replacement, a contracting relationship the state cannot size, and an absence of anyone measuring whether any of it produced more output.

5 · Research dependencies

Established Nothing in this brief waits on a research result. There is no discovery pending whose arrival would change the analysis. What civil service reform waits on is measurement, money and rules, and the absence of the first is the most consequential.

Established It depends first on an output measure that somebody outside the department chooses. Every favourable claim in this subject — savings, capability improvement, productivity from automation — is currently produced by the body being assessed or by a supplier selling the method, and every adverse finding comes from an audit office, a parliamentary committee, a royal commission or a statutorily independent review body. Those are not two interpretations of the same evidence; they are two different classes of number, and the second class exists only where an institution was built to produce it.

Established It depends on pay that is competitive for the specific skills that are scarce, which is not the same as higher pay. Uniform percentage settlements applied across a compressed structure widen the gap where the market gap is largest, because the same percentage is a smaller absolute sum at the top and in the highest-cost regions. The relevant dependency is a structural one, and the Institute for Fiscal Studies puts it precisely: government needs to structure remuneration to get the right people in the right roles, rather than simply raising the paybill.

Established It depends on tenure protections that are not removable by executive instrument. The comparative evidence associates merit protection with performance and with lower corruption; the theoretical work predicts that removing it drives out the most motivated first and selects for those willing to accept the new terms. That makes protection a dependency with a ratchet: cheap to remove, slow and possibly impossible to restore with the same people.

Frontier It depends, weakly and reversibly, on contracting relationships whose size is unknown. The dependency is real — day rates 61% above internal equivalents in the Canadian case, a knowledge-transfer process the United Kingdom auditor found to be systematically neglected — but it is a dependency the state entered by decision and could in principle exit. Whether it can in practice, having lost the staff, is the open question in the frontier section.

Established What depends on this is most of the rest of the map. Future Public Administration depends on it entirely, since a delivery record is a record of what a workforce delivered. AI-Assisted Governance depends on the digital-skills figures and on the quality-assurance capacity identified here as a bottleneck. Institutional Design supplies the finding that reorganisations rarely define a success measure, and this brief supplies four instances of it. And every brief on this map whose constraint is delivery rather than discovery depends on there being an administration capable of delivering — a dependency this brief declines to assert as an enabling edge, for reasons stated in the tech tree.

6 · Required experiments

Established The cheapest decisive experiment has already been designed twice and simply needs an output measure attached. Three United Kingdom departments evaluated the same product in the same quarter with three designs. Adding one more arm — random assignment of licences within a single processing function, with the dependent variable being cases cleared per week rather than minutes reported — would settle in one quarter the question the entire policy programme currently assumes. The administrative data already exist, because processing functions count their own throughput for other reasons. Nobody has run it.

Frontier The natural experiment already running is the United States federal workforce. A service that shrank 12% in sixteen months while its under-30 share fell and its AI use-case count doubled is the strongest test of the substitution hypothesis anyone will get, and the readouts are public: separations and accessions by occupation, processing times, and backlog series at agencies whose caseloads did not fall. Tax positions fell about a third and the IRS fell 25%; food inspection fell more than half; park rangers fell 34.1%. If administrative automation substitutes for administrative labour, service metrics at those agencies should be flat or improving. This brief takes no position on the outcome and records the design.

Frontier The second natural experiment is Schedule Policy/Career, and its virtue is that it has a partial treatment. Roughly 8,000 positions were moved against an estimate of 50,000, overwhelmingly at GS-15 and above, concentrated in particular agencies. That produces treated and untreated units within the same government at the same time, which is exactly the comparison the observational politicisation literature has always lacked. The outcome measures the literature already uses — non-competitive contract award rates, response times to information requests, voluntary exit rates among career staff — are administratively available. The design is available whether or not anyone runs it, and the window closes when the schedule is either extended or reversed.

Established A negative result already recorded and worth preserving as one. The DBT evaluation is a departmental study that looked for a productivity effect of a tool its own government was promoting, did not find one, and published that. It also published the observed-task data showing the tool making one task slower and less accurate. Negative results of this kind are rare in government evaluation because the incentive runs the other way, and this brief weights it more heavily than the two positive studies precisely because of the direction it points relative to the interest of the body that produced it.

Frontier The experiment that would settle the consultancy question is a records experiment, not an economic one. The National Audit Office's finding is that four data sources disagree by up to £890 million on one year because definitions differ and contracts bundle. Applying a single definition retrospectively to a sample of contracts across departments and arm's-length bodies would establish, for the first time, the size of the dependency — and would incidentally test whether the reported fall from £1.57bn in 2021-22 to £1.36bn in 2022-23 was a real reduction or a reclassification.

Established And the experiment nobody wants to run, which the Auditor General of Canada has effectively scheduled anyway. Replacing a government-wide pay system on a timeline shortened by three years, with 233,653 transactions still in backlog and 155,217 of them over a year old, is a live test of whether an institution can learn from a documented failure of exactly the same shape. The auditor has stated the risk in advance and in public, which means the result — whichever way it goes — will be interpretable rather than contested, and that is a rare property in this field.

7 · Engineering requirements

Established The instruments a civil service actually operates are five, and each has a measured failure mode. Headcount controls, which reverse across a spending round. Pay bands, which compress and then produce exit-and-return. Spend approvals, which accumulate until the centre is personally approving two hundred decisions a year worth less than a million pounds each and are then cut by 55% for diluting the accountability they were meant to enforce. Appointment and tenure rules, which generate complaint volumes the regulator cannot act on. And contracting, which is the one instrument whose own volume the state cannot measure.

Established The clearest way to see what an AI assistant does to administrative work is the observed-task table, and it belongs here rather than in the position section because it is the mechanism. These are directly measured task completions with blind-scored output, three per condition, from a departmental evaluation. Quality and accuracy are scored one to five.

TaskSelf-reported savingTime with assistantTime withoutAccuracy with / without
Summarising a report+0.8 hours12 min 37 s41 min 34 s4.0 / 2.5
Data analysis in a spreadsheet+0.6 hours25 min 1 s20 min 33 s1.5 / 2.7
Writing an email+0.2 hours7 min 30 s7 min 43 s4.3 / 4.0
Preparing a presentation≈ 0about 7 min faster1.5 / 5.0 (quality)
Schedulingnet loss of ≈ 0.6 hours

Established Three things follow from that table and none of them is “AI saves time”. First, the sign of the effect depends on the task, and the two tasks with the largest effects have opposite signs. Second, self-report does not track measurement in either direction — the task with the largest true gain and the task with a true loss produced similar self-reported savings. Third, on presentations the tool was faster and the output was scored 1.5 against 5.0 for quality, which is the shape of a productivity statistic that is really a quality statistic in disguise. The condition to state with every one of these numbers is the sample: three observed tasks per condition, in one department, over three months, on one product version, with the report's own caution that external factors cannot be excluded at that sample size.

Established Adoption behaves the way adoption always behaves, which is unevenly. In the cross-government experiment about 80% of licence-holders remained active, but usage split sharply by application — 71% in the meetings client against 23–24% in the spreadsheet and presentation tools, with usage in the word processor falling 5% and in the mail client 10% from their peaks. Departments and professions with the lowest prior familiarity and confidence saw the lowest benefit, which is the opposite of the distributional story usually told about these tools. In the Pennsylvania pilot, 30% became heavy users, 48% used the tool sporadically, 18% used it for one or two specific things, and 4% did not use it at all — the last figure being, on the report's own account, an undercount.

Established Hallucination is not an edge case in these deployments; it is a routine operating condition with an unmeasured quality-assurance response. 22% of diarists in the DBT evaluation reported encountering hallucinations, 43% reported none, and 11% did not know. More consequential: depending on the task, between 23% and 64% of users either did not quality-assure the output or could not compare it to an unaided baseline. In Pennsylvania the reported failure modes were fabricated citations and links, broken text extraction from PDFs, and legal answers that one user said “were almost never correct”. A deployment in which a fifth of users see fabrications and up to two-thirds do not check is a deployment whose error rate is unknown by construction.

Established The public-facing failure case has a price attached. New York City's MyCity business chatbot, launched in autumn 2023, told users that landlords could refuse housing vouchers, that employers could take a share of workers' tips, and that businesses could refuse cash — the last contrary to a 2020 city law. It cost nearly $600,000 to build and about half a million dollars a year to run. After the errors were documented by investigative reporters in March 2024, the administration added a disclaimer telling users not to treat its answers as advice and left it running for nearly two more years; it was shut down in February 2026 as “functionally unusable”. The interesting number is not the error rate but the twenty-two months, which is a measurement of institutional response time rather than of the technology.

8 · Adjacent technologies

Within this map the nearest neighbour is Future Public Administration, and the two briefs split by unit of analysis rather than by subject: this one owns the workforce and its rules, that one owns programmes and their delivery. They share the Phoenix and ArriveCAN material deliberately and take different halves of it — the pay backlog and the contractor day rates here, the project management and the procurement decisions there.

AI-Assisted Governance is adjacent in the strong sense that it is downstream of this brief's numbers. Any proposal for automated or assisted decision-making in government is priced by the digital-skills figures, by the fill rate, and above all by the quality-assurance capacity identified here as the binding constraint. The three evaluations summarised in this brief are the empirical floor under that one.

Institutional Design supplies the general finding that reorganisations rarely define a success measure, and this brief supplies four instances: a digital unit dissolved as “non-critical” with no evaluation cited, a second renamed and repurposed, a third relocated into a delivery department, and a fourth expanded by target. Scientific Advisory Institutions is adjacent because expert advisory capacity inside government is subject to the same tenure and demographic pressures described here.

Outside the map: personnel economics and the literature on public-sector pay compression, which supply the mechanism connecting pay structure to fill rates; industrial relations, where the morale and consent evidence properly belongs; the comparative politics of bureaucratic autonomy, which is where the 96-study merit literature sits; and administrative law, which is where automated decision-making meets the question Robodebt answered the hard way — whether a system that averages can lawfully decide.

Two relationships are worth naming precisely because they are usually assumed rather than argued. Procurement is adjacent in the strong sense: the contractor day rates, the intermediary firms and the irreconcilable consultancy figures in this brief are procurement facts that become workforce facts, because every pound spent on a contracted capability is a pound not spent building an internal one, and the auditor's finding on knowledge transfer is what turns that arithmetic into a ratchet. And the technology vendors are adjacent in a way the literature rarely acknowledges: the entire published evidence base on AI assistants in a civil service is an evidence base on one company's product, procured through enterprise agreements that predate any evaluation, which makes vendor selection an upstream determinant of what gets measured at all.

9 · Institutional requirements

Established The institution that decides the future of civil services is not a civil service. Every consequential decision recorded in this brief was taken by an executive instrument or a machinery-of-government reshuffle: a reduction in force dissolving a digital unit as “non-critical”, an executive order moving 8,000 posts out of merit protection with seven days for agencies to update their records, a renaming of an agency, a transfer of a unit from a central agency into a delivery department. None of them was preceded by a published evaluation, and none of them is reversible by the body affected.

Established The institutions that produce the reliable numbers are the supreme audit institutions, and they are the reason this brief can be written at all. The Auditor General of Canada established that ArriveCAN's cost was unknowable from the government's own records and priced the contractor premium at 61%; the same office established the Phoenix backlog at 233,653 transactions and put the Dayforce estimate at over $4.2 billion while flagging the three-year schedule compression as a risk in advance. The United Kingdom's National Audit Office established that four sources give four consultancy figures. The United States Government Accountability Office established that agency AI inventories are not comprehensive and that a majority of generative-AI use cases were not in operation. The selection bias is worth naming: the countries with the best-documented civil service failures are the countries with the best-functioning auditors, and the absence of a comparable record elsewhere is not evidence of a better record.

Frontier The institution that does not exist is an independent evaluator of administrative automation. Three departmental evaluations of the same product produced three different answers in the same quarter, using three designs, with no common protocol, no pre-registration, no shared outcome measure and no independent body reconciling them. The most rigorous of the three was produced by a department evaluating a tool its own government was promoting, which is a fortunate accident rather than an institutional guarantee. Nobody is required to measure whether an AI deployment in a government agency changes output, and consequently nobody has.

Frontier The regulator of appointments is narrower than the grievances brought to it, which is a design fact rather than evidence about appointments. Of 216 recruitment-principles complaints, 183 fell outside the regulator's legal remit; of 120 civil service code appeals, all 120 did. Senior competitions fell from 235 in 2023-24 to 166 in 2024-25, 32.5% produced only one appointable candidate, and 7.2% of vacancies went unfilled. This brief takes no position on whether appointments are becoming politicised in the United Kingdom, because no published exception totals or ministerial-involvement figures exist — and it notes that the United States, where the equivalent change was made by public instrument, is the more measurable case precisely because the change was explicit.

Established The buyer of the tools is an interested party and the seller is a monopolist by default. Microsoft Copilot appears in 102 reported United States federal AI use cases; the three United Kingdom evaluations are all evaluations of the same Microsoft product; the New York City chatbot ran on the same vendor's technology. The evidence base on AI in government is therefore substantially an evidence base on one company's product, procured through existing enterprise agreements, which is how it came to be deployed at 20,000 desks before anyone had measured whether it changed output.

Established And several bodies in this story no longer exist under the names commentary still uses. 18F does not exist. The United States Digital Service was renamed in January 2025. The Canadian Digital Service sits inside Employment and Social Development Canada rather than the Treasury Board Secretariat, where a co-founder observed that what is lost is the whole-of-government perspective and a voice at the cabinet table. The United Kingdom's central digital and data office was merged into a single digital service inside a science department in January 2025. Any assertion about a named central digital body should carry a date, and most published assertions do not.

10 · Ethical & societal considerations

Established The standing evidential asymmetry in this subject is the first ethical fact and it should be stated before any number is read. Every favourable claim — savings, capability improvement, productivity from automation — originates with the body being assessed or with a supplier selling the method. Every adverse finding originates with an audit office, a parliamentary committee, a royal commission or a statutorily independent review body. The two groups are not disagreeing about interpretation; they are producing different classes of number, and the second class exists only where an institution was deliberately built to produce it. This brief weights audit findings above departmental claims as a rule, and marks interested parties in its reading list.

Established The single most credible source in the AI section is credible because of the direction it points relative to its own interest. A departmental evaluation that looked for a productivity effect of a tool its own government was actively promoting, did not find one, published that finding in plain language, and printed the observed-task data showing the tool making a task slower and less accurate — that is a document produced against interest, and it is weighted here above two studies that found larger effects with weaker designs.

Frontier There is a specific ethical problem in deploying tools whose error rate is unknown by construction. Between 23% and 64% of users, depending on the task, did not quality-assure output or could not compare it; 22% encountered fabrications. In a commercial setting that is a business risk borne by the firm. In an administration it is borne by whoever the decision is about, and the New York City chatbot shows the shape: a public system telling businesses they could discriminate against voucher holders and take a share of workers' tips, left running for nearly two years after the errors were documented, with a disclaimer added telling users not to rely on it. A disclaimer transfers liability without reducing harm, and it was the institution's chosen response.

Established Robodebt is the case that establishes the stakes and it should not be softened. An automated scheme that averaged annual income across fortnightly reporting periods produced, in the Royal Commission's finding, results that were inaccurate and did not comply with the income calculation provisions of the governing Act. The Commissioner's assessment of the institution rather than the software is the part that belongs in this brief: “It is remarkable how little interest there seems to have been in ensuring the Scheme's legality, how rushed its implementation was.” Its recommendation categories read as a specification for what a civil service is for — mandatory legal risk assessment in policy proposals, oversight of automated decision-making, strengthened duties on government lawyers, better-resourced oversight agencies.

Frontier There is an ethical dimension to workforce reduction by voluntary exit that is not usually posed as one. A deferred-resignation programme is voluntary at the level of the individual and coercive in aggregate, and its composition effect — older, more experienced staff leaving while the entrance is closed — transfers a cost from the present budget to a future capability that nobody is accountable for. The published record contains no evaluation of what a service lost when 136,822 people took such an offer, and the brief that this one replaces made the same observation about a headcount cut that was fully reversed within nine years: a reduction that does not survive, or that removes the wrong people, has extracted the human cost of a reform without delivering the reform.

Established Finally, an obligation this brief takes on itself, stated as unresolved questions rather than as caveats. It cannot say what share of central government employees across the OECD is aged 55 or over, because the comparative source would not open to it and the demographic material here is therefore American and Canadian rather than international. It has not read the 96-study meta-analysis on which its politicisation argument rests; that finding is reported at second hand through a university research centre's summary and is named as such. It has not read the Robodebt final report itself, only a government agency's published overview of it, and it therefore carries none of the figures on debts raised or settlement value that are usually quoted. It has verified no Estonian or Singaporean figure at all, which means the two exemplars most often cited in this field are absent here rather than assessed. And it cannot say whether the historical computerisation of government substituted for clerical labour, because the one study it found says the opposite of the common intuition and it could read only the abstract. Each of those is a claim this brief could have made and has not.

11 · Civilizational implications

Established The civilisational content of this subject is not the size of the state; it is whether a society can maintain a body of people who know how it works. Everything measured in this brief is a measurement of that stock. Twenty times as many information-technology staff over 50 as under 30 is a statement about institutional memory with a fixed expiry date. Contractor day rates 61% above the internal equivalent, on work whose knowledge transfer the auditor found to be systematically neglected, is a statement about where the knowledge now lives. Eight thousand posts moved out of merit protection is a statement about the terms on which it is held. These are not three subjects; they are three readings of one gauge.

Established The recurring general principle is that reform apparatus measures itself. Capability scores improved with no demonstrable link to delivery, on measures the assessing committee called qualitative and subjective. Savings were computed from data used for no other business purpose. Reform commitments were delivered as announcements at a rate of about one in four. And now AI evaluations measure minutes of self-reported time on task rather than cases cleared. Four decades apart, the same structural error: the metric is generated by the party with an interest in it, and the independent number, where one exists, moves the other way.

Frontier Where automation becomes civilisationally significant, the mechanism is probably not substitution. The best direct evidence says the tool is strong at compressing text and weak at analysis, which points at a change in the composition of administrative work rather than its volume — less drafting, no less judgement, and more of a new activity that did not exist before: checking machine output. The binding constraint on that new activity is the number of people with enough domain knowledge to know when the machine is wrong, which is exactly the population that ageing, pay compression and contracting-out are removing. If that is right, automation and the demographic transition interact badly rather than solving each other, which is the opposite of the assumption on which most current policy rests.

Speculative The larger question is whether a permanent professional bureaucracy is a stable institutional form or a historically contingent one. It is roughly 150 years old in most countries that have one, it was built to solve a specific problem — patronage — and the evidence that it solved that problem is as good as evidence in this field gets. What is being tested now, in real time and without a control group, is whether it survives simultaneous pressure from automation, from fiscal contraction applied through voluntary exit, and from an explicit argument that tenure protection is itself the pathology. This brief does not forecast the outcome. It records that the test is running and that the instruments to measure it exist and are not being used.

12 · Timelines

Established What already happened, because the chronology usually starts too late and stops too early. 18F was founded in March 2014 and terminated in March 2025; the United States Digital Service was renamed in January 2025 and had lost about a third of its staff by mid-February; Canada's CDS was moved out of its central agency in July 2023. Phoenix went live in 2016 and its backlog stood at 233,653 transactions in September 2025. Australia's Robodebt scheme ran from 2015, stopped raising averaged debts in November 2019, settled in 2020 and was reported on in July 2023. The Auditor General of Canada reported on ArriveCAN in February 2024. None of these is a pending milestone; all are completed events that most writing on the future of civil services does not reference.

Established 2024 to 2026: the evaluation window that has now closed. The three United Kingdom Copilot studies all ran between October 2024 and March 2025 and reported between June 2025 and February 2026. That is the entire published evidence base on general-purpose AI assistants in a civil service, and it is already historical: the products have changed version several times since, and the DBT evaluation notes that features introduced during its own trial were outside its scope. Any figure quoted from these studies is a measurement of a 2024 product.

Frontier 2026 to 2027: the targets that will be marked. The National Digital Exchange is to deliver £1.2 billion a year in savings and new funding models launch in summer 2026; consultancy spend is to fall by over £700 million a year by 2028-29 from a baseline that four sources cannot agree on; and the Dayforce planning phase is to complete by June 2027. The first two are marked against baselines the government's own auditor says are unmeasurable, which is a prediction about what the marking will find.

Handwave 2030: one civil servant in ten in a digital or technology role. The March 2025 actual is 5.2%. Reaching the target requires roughly doubling a profession whose recruitment fill rate has already halved, in a labour market where the state pays furthest below the private comparator for exactly these skills, without a stated mechanism for the pay problem. This is recorded as an intention. Targets in this subject have a long record of being restated rather than met, and the published brief's own scorecard — roughly one commitment in four delivered as an announcement rather than a change — is the base rate.

Frontier 2031 to 2036: the pay system. Dayforce was scheduled to replace Phoenix by 2034 and in January 2026 the timeline was shortened by three years; Phoenix vendor support runs to as early as 2036, with cloud extension costing at least $4 million a year in the meantime. The preliminary cost estimate is over $4.2 billion and excludes transition. The auditor's three flagged risks — schedule compression, an uncleared backlog importing errors, and unsimplified pay rules forcing customisation — are the three causes of the failure being replaced.

Speculative Late 2020s: the demographic transition completes whether or not anything is done about it. A technical workforce with twenty times as many people over 50 as under 30 resolves itself within a decade by arithmetic. What replaces it — a smaller cadre with better tools, a contracted-out function, or a gap — is not forecastable from the current record, and the three outcomes are not distinguishable in advance by any indicator this brief could identify.

13 · Technology tree & dependencies

  • Depends on Nothing on this map. No research result is pending whose arrival would change what a civil service can do. The constraints are age structure, pay structure, tenure rules, contracting relationships and the absence of an independent output measure, and every one of them is recorded below rather than being awaited from another brief.
  • Requires (not on this map) An output measure chosen from outside the department: on the record here, capability programmes improve their own scores, savings figures are computed from data used for no other business purpose, and AI evaluations measure minutes rather than cases. Pay structured for the skills that are actually scarce rather than raised uniformly — the compression evidence shows uniform settlements widening the gap exactly where the market gap is largest. Tenure protection that an executive instrument cannot remove, since roughly 8,000 positions moved out of the merit system by order in a single year against a literature associating merit protection with performance across 96 studies and 150 countries. Enough domain experts to quality-assure automated output, which is the binding constraint on AI deployment and the same population the age structure is removing. And contracting records good enough to audit: one auditor could not determine what a single application cost, and another found four irreconcilable figures for one year's consultancy spend. All five are institutional or industrial capabilities, not discoveries.
  • Enables In principle, administrative capacity for every brief on this map whose bottleneck is delivery rather than science. No typed enabling edge is claimed, and the reason is the brief's own standard of evidence: the reform programmes examined here improved their own metrics without a demonstrable link to delivery, and the AI deployments measured time on task without measuring output. Asserting an enabling edge would assert exactly the causal claim this brief says nobody has measured.
  • Adjacent Future Public Administration, which owns the delivery half of the same subject; AI-Assisted Governance, whose feasibility is priced by the digital-skills and quality-assurance figures here; Institutional Design, for the finding that reorganisations rarely define a success measure; Scientific Advisory Institutions; and outside the map, personnel economics, the public-sector pay literature, industrial relations, and the comparative politics of bureaucratic autonomy.

14 · Common misconceptions & speculative claims

“AI saves civil servants about 26 minutes a day.” Frontier The 26-minute figure is a survey answer from a 20,000-person trial with no comparison group, whose own limitations section warns against taking self-reported time savings at face value and states that it could not identify how the saved time was spent. The same product measured against a comparison group at another department gave 19 minutes. The same product evaluated with directly observed, blind-scored tasks at a third produced no evidence of a productivity improvement at all, with users measurably slower and less accurate at spreadsheet analysis. The three results are not in conflict; they are a dose–response curve in study design. Quote the range and the designs, or quote nothing.

“Pennsylvania proved generative AI saves 95 minutes a day.” Handwave One hundred and thirty-six exit-survey respondents from 175 volunteers across 14 agencies, self-reporting, in a sample the report itself says is not representative, with an acknowledged undercount of people who stopped using the tool and an admitted shortfall against its own pre-pilot expectation. It is a useful pilot report and a bad statistic, and the report's own concluding sentence — that the tool is not a substitute for the nuance and experience employees bring — is the part that does not travel.

“Agencies are deploying thousands of AI systems.” Frontier The count is real and the interpretation is not. 3,611 reported United States federal use cases in 2025 is a reporting artefact: in the auditor's examination of the prior year, 30% of generative-AI use cases were merely initiated and 27% in acquisition or development, a majority not in operation, and 61% were back-office mission-enabling functions rather than services to the public. In an earlier review, only 5 of 20 agency inventories were comprehensive, with duplicates and non-AI technology included, and one agency reported 375 uses to the auditor against 33 published. The Pentagon and Intelligence Community do not report at all. A use-case count measures compliance, not deployment.

“The digital service units failed and were shut down.” Established No evaluation is cited in any closure examined here. 18F built Login.gov, cloud.gov, the US Web Design System and COVIDtests.gov, ran for eleven years on a cost-recovery model that required client agencies to choose to pay for it — about as close to a market test as a public body gets — and was terminated by email at 1 a.m. as “non-critical”. The United States Digital Service was renamed and its staff cut by a third, after which twenty-one technologists resigned rather than be redirected. Canada's CDS was moved from a central agency into a delivery department. The unit that survived is being expanded by prime-ministerial target. These are decisions about org charts and executive authority, and the performance question was not asked in any of them.

“The civil service is ageing and needs to recruit more young people.” Established True and beside the point, because the standard instruments make it worse. A deferred-resignation programme plus a hiring freeze removes people who were closest to leaving and closes the only channel through which a workforce gets younger: the United States federal under-30 share fell from 8.9% to 7.9% during a year in which the workforce shrank 12%. The binding constraint is not the size of the recruitment target but that the entrance was closed while the exit was widened, and that in the technical professions the succession has already failed — twenty times as many information-technology staff over 50 as under 30.

“Schedule F will politicise 50,000 federal jobs.” Frontier About 8,000 positions had actually been moved as of June 2026, roughly 97% at GS-15 and above, against an Office of Personnel Management working estimate of 50,000 and earlier projections reaching 200,000. The gap is the story in both directions, and this brief declines to resolve it: implementation is far short of the announcement, and the same instrument that moved 8,000 can move the rest without further authority. Litigation is live and its outcome is not forecast here.

“Merit versus political appointment is a values argument.” Established It has a quantitative literature. A meta-analysis of 96 peer-reviewed studies across more than 150 countries associates merit recruitment, promotion and tenure protection with higher government performance and lower corruption; a separate 52-country review reaches the same association; a one-standard-deviation increase in politically appointed managers is associated with a 51–63% higher probability of non-competitive contract award; appointees average 18–24 months in post; and there is no evidence of partisan cycling among career staff against significant cycling among appointees. The evidence is correlational and must be described as correlational — countries with merit systems tend to have other good institutions for reasons predating both — but it is not absent and it does not point both ways.

“ArriveCAN cost $59.5 million.” Established That is the auditor's estimate, produced because the records could not answer the question: 18% of tested contractor invoices did not document which project they belonged to, and the audit's own words are “we were unable to determine a precise cost”. Quoting $59.5 million as a known figure inverts the finding, which is about the absence of records rather than the size of a bill. The findings that survive precisely are the day rates — $1,090 external against $675 internal — and the 177 application releases with little or no documentation of testing.

“Government spends £X billion a year on consultants.” Established No such figure exists. Four sources give £1.34bn, £1.36bn, £1.68bn and £2.23bn for the same year, with an average variance of £270 million a year over six years, because definitions differ, contracts bundle consultancy with professional services and contingent labour, and arm's-length bodies report separately. Any single figure quoted without a definition is one of four incompatible numbers, and a reported year-on-year fall may be a reclassification.

“Phoenix is being fixed.” Frontier 233,653 pay transactions were in backlog at 30 September 2025 with 155,217 of them more than a year old, affecting 133,619 employees — a 27% reduction on April 2023, which is the measure of how slow the recovery has been on a system paying $38 billion a year to more than 430,000 people. The replacement is preliminarily estimated at over $4.2 billion excluding transition, and its schedule was shortened by three years in January 2026 — the same class of decision the auditor identifies as the Phoenix risk.

“Computerisation of government has historically just added a layer.” Handwave This brief will not assert it. The one empirical study located points the other way: state government IT investment was found to have positive and significant effects on gross state product, with larger returns where a chief information officer structure existed. One study, one country, one tier, published in 2002, and only its abstract was accessible. The sceptical intuition is currently an intuition, and saying so is the difference between scepticism and a prior.