1 · Concept overview

Three founding texts, three different claims, routinely merged into one. W. Ross Ashby coined “intelligence amplification” in An Introduction to Cybernetics in 1956, and meant by it any mechanism that improves the ratio of appropriate selection — no computer required. J. C. R. Licklider’s “Man-Computer Symbiosis” appeared in IRE Transactions on Human Factors in Electronics HFE-1, pages 4–11, in March 1960, and proposed a division of labour between a human who sets goals and formulates hypotheses and a machine that does the clerical work in between. Douglas Engelbart’s “Augmenting Human Intellect: A Conceptual Framework”, Stanford Research Institute, October 1962, proposed something larger and stranger: a co-evolving system of a human, the language he thinks in, the artifacts he uses, the methodology he follows, and the training that binds them. On 9 December 1968 Engelbart put ninety minutes of that system in front of roughly a thousand computer professionals at the Fall Joint Computer Conference in San Francisco, and the industry took the mouse.

This brief argues three things. First, that the augmentation tradition’s central empirical promise has now been tested at scale exactly once, and the result is negative: human-AI combinations average below the better of their two components, and that is the finding that survives a publication-bias check while the cheerful companion finding does not. Second, that the tradition’s most-repeated exemplar — centaur chess — has no empirical literature whatsoever, inconsistent protagonists’ names, and unsourceable ratings, and that this is a serious fact about how the field reasons. Third, that Engelbart’s actual programme was never tried: three of his four classes of augmentation means were not technological, the commercial descendants kept only the fourth, and the experiment that would settle whether the other three were load-bearing has never been run. The question the brief ends on is whether augmentation is a stable equilibrium or a moving boundary that dissolves from below as machines improve. The honest answer is that one cheap experiment would tell us and nobody has run it.

2 · Current scientific position

Established The term is Ashby’s, it is 1956, and it is not a claim about computers. Ashby’s argument in An Introduction to Cybernetics is that problem-solving is appropriate selection from a space of options, and that if the power of selection can be amplified, it seems to follow that intellectual power, like physical power, can be amplified. Established Nothing in that formulation privileges a machine. Speculative On Ashby’s own terms the largest intelligence amplifiers in history are social technologies — double-entry bookkeeping, peer review, the surgical checklist, the market price — and the computer-centred tradition that inherited his phrase is one narrow and possibly unrepresentative branch. Frontier No common metric exists on which to test that branch against the computational one, so it appears here as a live hypothesis rather than a finding.

Established Licklider’s 1960 paper is a division-of-labour argument grounded in a self-experiment, and its structure matters more than its slogan. Part II of “Man-Computer Symbiosis” sets out the aims of symbiosis; Part III is a preliminary and informal time-and-motion analysis of technical thinking, built from a log Licklider kept of his own working activities in 1957. Frontier The famous percentage attached to that self-study — the share of his thinking time that went on clerical operations rather than on decisions — is quoted everywhere, could not be verified for this brief, and is therefore not printed here. Speculative A second and more consequential claim about the same paper is that Licklider framed symbiosis as an interim arrangement, valuable only until machines outdo the human brain in the functions he describes. Frontier The structure of that claim is well attested in secondary accounts and it is the historical anchor of the waypoint hypothesis this brief returns to in sections 3 and 14. Handwave It is also unverified at the level of wording and interval, and this brief names it as an open verification task rather than a correction: someone needs to read pages 4 to 11 of HFE-1 and report what Licklider actually wrote. If it holds, the founding paper of the augmentation tradition anticipates the tradition’s own obsolescence, which would be the single most useful fact in this subject.

Established Engelbart’s 1962 framework states an objective that is not about speed and not about replacing judgment. The stated goal is increasing the capability of a man to approach a complex problem situation, to gain comprehension to suit his particular needs, and to derive solutions to problems — wording reproduced identically in two independent sources consulted for this brief. Established Note what it does not say: no throughput claim, no substitution claim, and the phrase about comprehension suiting the user’s particular needs is a claim about the human’s model of the problem rather than the tool’s model of the human. Frontier The report, prepared for the Air Force Office of Scientific Research, carries a number is given two ways by sources of equal authority — AFOSR-3223 and AFOSR-3233 — and this brief cites the commoner form while noting that the discrepancy is unresolved. Established The load-bearing structure is the H-LAM/T system: a Human using Language, Artifacts and Methodology, in which he is Trained. Frontier Engelbart’s own wording of those four classes of augmentation means could not be obtained for this brief, and his weighting of them — which decides a central question in section 14 — is reported as attested rather than as read. Established The arithmetic is not in doubt: three of the four means are not technology.

Established The 1968 demonstration was a distributed-collaboration demo, and the retellings drop the collaboration. On 9 December 1968, at the Fall Joint Computer Conference in the San Francisco Civic Auditorium, Engelbart presented for ninety minutes to roughly a thousand computer professionals. The itemised list of what was shown is: windows, hypertext, computer graphics, efficient navigation and command input, video conferencing, the computer mouse, word processing, dynamic file linking, revision control, and a collaborative real-time editor. Established The physical stack is part of the argument: an Eidophor projector onto a 6.7-metre screen, two custom homemade 1200-baud modems over a leased line from the auditorium to the SDS 940 at Menlo Park, and two microwave links carrying live two-way video between the lab and the hall, with Bill English running a video switcher and Stewart Brand on camera thirty miles away. Funding came from NASA and ARPA; the team included Bill English on technical direction, Bill Paxton and Jeff Rulifson. Frontier The two video links exist so that the audience can watch Engelbart work with a colleague in another building, which is the thesis.

Established The programme died of budget first and of pedagogy second, and both causes are documented. NLS ran on a CDC 160A in 1963, a CDC 3100 in 1965, the SDS 940 for the demo in 1968, and a PDP-10 from 1970, funded by ARPA, NASA and the US Air Force. Established The 1969 Mansfield Amendment ended military funding of non-military research; the end of Vietnam and the termination of Apollo compounded it, gradually draining the Augmentation Research Center’s funding through the early 1970s. Established SRI placed the lab under an artificial-intelligence researcher, Bertram Raphael, who negotiated its transfer to Tymshare in 1976 or 1977 depending on which source is read; NLS was renamed Augment, and Tymshare was bought by McDonnell Douglas in 1984. Established The separate, pedagogical cause is stated plainly in the record: NLS had a difficult learning curve, heavy use of program modes, a strict hierarchical structure, no point-and-click interface, cryptic mnemonic codes and a chord keyset with a 5-bit binary code. Established Engelbart prioritised making the user more powerful over making the system easier to use, and coined WYSIAYG — What You See Is All You Get — as the deliberate inversion of the industry’s slogan. Frontier Engelbart did not abandon the programme; everyone else did. Its later institutional vehicle was the Bootstrap Institute, founded in 1988 with his daughter Christina, which ran three-day and half-day management seminars at Stanford from 1989 to 2000. Frontier A research programme whose successor organisation is a management seminar has not been refuted; it has been defunded, which is a different thing and calls for a different response.

Established The augmentation thesis now has one quantitative test at scale, and it is negative. Vaccaro, Almaatouq and Malone, “When combinations of humans and AI are useful”, Nature Human Behaviour 8:2293–2303 (2024), is a preregistered systematic review and meta-analysis of 106 experimental studies and 370 effect sizes, drawn from the ACM Digital Library, Web of Science and the AIS eLibrary, covering 1 January 2020 to 30 June 2023. Established Inclusion required an original human-participants experiment reporting all three arms: humans alone, AI alone, and the combination. Established The headline: against the better of the two components, the pooled effect was g = -0.23 (t92 = -2.89; two-tailed P = 0.005; 95% CI -0.39 to -0.07). Established Against humans alone, the same combinations were clearly positive: g = 0.64 (t98 = 11.87; P < 0.001; 95% CI 0.53 to 0.74). Established Both are true at once, and reporting either without the other misleads. Established The publication-bias tests are the part almost nobody quotes, and they cut against the enthusiast twice. For the primary comparison against best-of-either — the negative result — Egger’s regression found no evidence of bias (beta = -0.67; t104 = -0.78; two-tailed P = 0.438; 95% CI -2.39 to 1.04), nor did the rank correlation test (tau = 0.05; P = 0.121). Established For the comparison against humans alone — the positive result — Egger’s regression does indicate bias (beta = 1.96; t104 = 3.24; P = 0.002; 95% CI 0.76 to 3.16). Frontier The pessimistic finding is the robust one and the optimistic one is inflated. This brief has not seen that stated anywhere else, and it should discipline every claim made about augmentation from here on.

Established The moderators are where the paper stops being a verdict and becomes a map. Decision tasks: g = -0.27 (t104 = -3.20; P = 0.002; 95% CI -0.44 to -0.10). Creation tasks: g = +0.19 (t104 = 1.35; P = 0.180; 95% CI -0.09 to 0.48). Established The creation confidence interval spans zero: what is established is that the two task types differ, not that combinations beat the better component on creative work. Established The sharpest result is the relative-baseline moderator. When humans outperformed the AI alone, the combination gained: g = 0.46 (t104 = 5.06; P < 0.001; 95% CI 0.28 to 0.66). When the AI outperformed humans alone, the combination lost: g = -0.54 (t104 = -6.20; P < 0.001; 95% CI -0.71 to -0.37). Frontier Read plainly: the combination helps while the human is the better component and hurts once the machine is. That is the empirical core of the argument that augmentation is a waypoint rather than an equilibrium, and it is the strongest single datum in this brief.

Established The counter-cases are real, well-run, and share one structure: they augment the weaker party toward an existing standard. PRAIM, the German nationwide mammography implementation reported in Nature Medicine in 2025, covered 463,094 women across 12 sites, 260,739 of them screened with AI support: detection 6.7 per 1,000 with AI against 5.7 without, a relative increase of 17.6% (95% CI +5.7% to +30.8%), with recall falling from 38.3 to 37.4 per 1,000. Frontier It is observational rather than randomised and carries no AI-alone arm, so it is at once the strongest existing evidence for complementarity in a decision task and inadmissible to the meta-analysis. Established Tschandl and colleagues (2020) found good-quality AI support in skin-cancer recognition improving accuracy over either AI or physicians alone, with the least experienced gaining most — and, in the same paper, that faulty AI can mislead the entire spectrum of clinicians, including experts. Established Tutor CoPilot, a randomised trial of 900 tutors and 1,800 K-12 students, raised topic mastery 4 percentage points overall (p < 0.01) and 9 points for students of the lowest-rated tutors, at about twenty dollars per tutor per year. Established Brynjolfsson, Li and Raymond’s 5,172 customer-support agents resolved 15% more issues per hour, the gain concentrated in the least experienced while the highest-skilled saw small speed gains and small quality declines. Frontier Every robust positive in this literature lifts the bottom of a distribution and every robust negative occurs at the top. That generalisation rests on five independent sources and is the most useful thing the augmentation tradition can currently claim — but it is a claim about variance reduction, not about raising a ceiling, and Engelbart’s programme was explicitly about the ceiling.

3 · Frontier questions

Frontier The hardest negative result lands squarely on the tradition and it is about experts, not novices. METR’s 2025 randomised trial gave 16 experienced open-source developers 246 tasks in mature repositories they had worked on for an average of five years, randomising whether early-2025 AI tools were allowed. Frontier The developers forecast a 24% reduction in completion time; after the study they estimated a 20% reduction; the measured effect was a 19% increase. Frontier Economists asked in advance predicted 39% faster and machine-learning experts 38% faster. Established The sample is 16 developers and should not be over-generalised, though 246 tasks is not small and the study is unusual in measuring practitioners on their own codebases. Frontier The finding inside the finding is the self-report gap: practitioners believed they were 20% faster while being 19% slower. Speculative If that gap generalises — and it replicates a very old human-factors result about the limits of introspection on automation-assisted performance — then most organisational evidence for augmentation, which is survey evidence, is measuring belief rather than output.

Frontier Set METR beside the customer-support result and the field’s central open question appears. Brynjolfsson and colleagues show augmentation compressing a skill distribution upward; METR shows it going negative at the top of one, on tasks where context is deep and tacit. Speculative The obvious reconciliation is that the sign of the effect flips with the ratio of tacit context to codified knowledge in the task. Handwave That is a hypothesis with no direct test behind it, stated here because it is the cheapest thing in this brief to falsify and nobody has tried. Frontier Peng and colleagues’ Copilot trial sharpens it: developers building an HTTP server in JavaScript from scratch finished 55.8% faster. Same intervention class, opposite sign, moderated by task context rather than model quality.

Frontier The most under-discussed frontier result for this subject is homogenisation. Ju and Aral’s field experiments with 2,234 participants producing 11,024 advertisements found human-AI teams delivering roughly 50% more output per worker with higher text quality, human-human teams producing higher image quality, and human-AI outputs that were measurably more homogeneous, or self-similar. Frontier The interaction mechanisms are instrumented for the first time: 17% more delegation to the AI than to a human partner, and 62% fewer direct text edits. Speculative If augmented individuals converge on similar outputs, augmentation raises the floor while compressing exactly the variance that collective problem-solving depends on, and no one has measured that at population scale. Handwave The strong version — that in search-like domains the diversity loss can exceed the aggregate individual gain — is a coherent argument with no measurement behind it at all.

Frontier The explainability escape route has been tested and it did not work. Bansal and colleagues (CHI 2021) ran mixed-method studies with an AI of human-comparable accuracy and found that explanations increased the chance that humans accept the AI’s recommendation regardless of its correctness, without improving complementary team performance. Frontier It is the most important negative result in the explainability literature and is regularly cited as though it said the opposite. Speculative It also suggests the deficit measured by the meta-analysis is not obviously an interface problem, which is the assumption most augmentation product work runs on.

Frontier The augmentation designs that look most like Engelbart’s are the ones where the pedagogy, not the model, does the work. Kestin and colleagues (Scientific Reports, 2025) ran a crossover trial with 194 Harvard introductory-physics students in which a custom tutor was informed by the same pedagogical best practices as the in-class lessons: linear-regression effect size 0.63, quantile-regression estimates 0.73 to 1.3 standard deviations, z = -5.6, and median time on task 49 minutes against about 60 in class. Frontier The authors declare no competing interests; no funding source is stated in the paper. Speculative In Engelbart’s vocabulary this is the methodology class of augmentation means doing the work with the artifact merely as carrier — which is the tradition’s own prediction about where the gains should come from, and it is being confirmed by people who are not citing him.

Frontier Two live concepts in this area rest on citations this brief could not verify, and that is itself a frontier fact. The “jagged frontier” framing, which organises much current thinking about which tasks augmentation helps, originates in a 2023 working paper whose bibliographic record could not be retrieved for this brief at all; its widely repeated headline numbers are not stated here. Frontier And the scholarly literature on centaur chess — the tradition’s own central case study — is empty: a bibliographic search returned zero empirical studies, with every hit on the term metaphorical. Established That absence is citable and it is discussed at length in section 14.

4 · Technological bottlenecks

Frontier The target this brief works back from is deliberately harder than anything the field currently attempts: a demonstrated, replicated co-evolutionary augmentation system whose human-plus-tool ceiling exceeds the best available automated system on the same task class, with the margin widening as the automated baseline improves. Established That is strictly harder than beating the human, which is routine, and strictly harder than beating the machine once, which is rare. Frontier Stated this way the target can fail, which is the point.

Frontier The first genuinely binding constraint is ex-ante identification of the human contribution. Complementarity that can only be recognised after the outcome is known is not deployable. A working augmentation architecture requires a policy, learnable from data and using no post-hoc outcome information, that predicts before the decision which instances the human should own — and the routed system must beat both arms. Speculative If human contribution is real but unroutable, then no deployable augmentation architecture exists and the strong structural pessimism about human-AI teams is true in practice whether or not it is true in principle. Frontier Everything upstream of this link is measurement; this is the first place the target can be impossible.

Frontier The second binding constraint is survival across a capability generation. Every existing complementarity result is a single-generation snapshot. Nobody has published the same task, the same routing policy retrained, with the automated baseline upgraded one full model generation, and reported whether the margin shrank. Frontier That series is the decisive test of whether augmentation is an equilibrium or a moving boundary, and the relative-baseline moderator — g = 0.46 when the human is better, g = -0.54 when the machine is — predicts that it shrinks. Established The experiment costs one re-run.

Frontier The third constraint is institutional rather than scientific, and it is the one that killed Engelbart. Demonstrating a co-evolution premium requires measuring a high-training-cost tool against a low-training-cost tool over enough months for the training cost to amortise. Speculative No product organisation runs on an eighteen-month evaluation horizon and few funders will support one, so the experiment that would settle whether Engelbart’s capability ceiling was real has never been run. Frontier The obstacle is funding structure, not method.

Established The fourth constraint is a measurement practice, and it is embarrassing. The meta-analysis could use only 106 studies from a far larger literature because most published work on AI assistance omits the AI-alone arm. Frontier A field that does not routinely measure the machine on its own cannot know whether its human-in-the-loop system is helping or costing, and has been reporting the wrong comparison for a decade. Established Fixing it requires no new technology and no new budget line; it requires a reporting standard.

5 · Research dependencies

Frontier This brief waits, first, on the teaming evidence base assembled in Human-AI Integration. The anchor meta-analysis, the automation-dependence record, the handoff-failure cases and the clinical reading studies are all developed there; this brief inherits its central number from that work and would have to be rewritten if the moderator structure moved. Frontier In particular, if the per-task-class map turns out to have classes with durable positive effects, the waypoint reading weakens and Engelbart’s ceiling claim gets a foothold.

Frontier It depends second on the collective-intelligence literature, because of homogenisation. Whether population-scale augmentation is a net gain or a net loss turns on how much collective performance depends on output diversity, and that is not this brief’s question to answer; see Collective Intelligence. Speculative The measurement exists on the individual-team side and does not exist on the population side.

Frontier Third, on the education evidence. The methodology and training classes of Engelbart’s framework are, operationally, education, and the standards of evidence there — instrument alignment, test-alignment collapse, publication bias — are the ones this subject needs; see Future Education Systems. Established Any augmentation effect size reported without naming the instrument that produced it is not a number.

Established Fourth, on archival access, which is a real and current constraint. The research behind this brief could not obtain Engelbart’s 1962 report, Licklider’s full text, or Kasparov’s 2010 essay on freestyle chess. Frontier Three of the sharper claims in sections 3 and 14 are held at the level of open verification tasks for exactly that reason, and the retrievals are individually trivial for anyone with library access.

6 · Required experiments

Established The chain is seven links and only two of them bind. Link 1: publish an operational taxonomy that classifies a deployment as augmentation or substitution on observable features — who holds the decision right, whose information enters the decision, what the reversion mode is — and demonstrate that two independent coders agree at kappa above 0.8 on a corpus of 100 deployments. Frontier The anchor meta-analysis’s 106-study corpus is publicly specified and is the obvious test bed. Cheap, undone, and every downstream measurement depends on it.

Frontier Link 2: extend the moderator analysis from two task classes to a pre-registered taxonomy of at least eight, with pooled effect sizes per class. The 370 effect sizes already exist. Frontier Link 3: demonstrate positive complementarity in a single decision task at p < 0.01 in a pre-registered replication, with the mechanism identified — and with the AI-alone arm included, which is the requirement most of this literature fails. PRAIM is the closest existing candidate and does not meet the bar, being observational and lacking that arm.

Link 4 binds: show that the human contribution is identifiable in advance. Build a routing policy from held-out data, using no post-hoc outcome information, and show the routed system beating both arms. Speculative The useful preliminary is the oracle-router bound: compute what perfect ex-post routing would have achieved on an existing dataset before claiming a deployable complementarity. If that bound is small, the argument is over cheaply.

Frontier Link 5 binds, and it is the experiment that decides this brief’s central question: re-run Link 4 with the automated baseline advanced one model generation, the routing policy retrained, everything else held. If the complementarity margin holds, augmentation is a stable architecture; if it shrinks, it is a waypoint. Established Nobody has published a two-generation series. Frontier It should be run in three unrelated domains before anyone believes the answer.

Speculative Link 6 is the Engelbart experiment and it has never been attempted: two matched cohorts on one task, one given a low-training-cost tool and one a high-training-cost tool with a deliberately higher designed ceiling, measured at 1, 6 and 18 months. The co-evolution cohort must cross over and stay ahead; the reportable quantities are the crossover point and the post-amortisation gap. Frontier Link 7 is bootstrapping proper: a group applying the Link 6 stack to improving its own improvement process, against a control that improves only object-level work, with a second-derivative estimate over at least three measurement periods. Handwave That is the untested core of what Engelbart actually claimed, and it cannot even be assessed until Link 6 exists.

Established Three archival retrievals belong on the same list because they are cheaper than any of the above. Read pages 4 to 11 of IRE Transactions on Human Factors in Electronics HFE-1 and settle whether Licklider framed symbiosis as temporary. Established Read the 1962 SRI report and settle whether Engelbart weights artifacts first or last among his four means. Frontier And run one measured centaur-versus-engine match post-2020, which would settle a claim carrying enormous rhetorical load and appears never to have been attempted.

7 · Engineering requirements

Frontier The engineering requirement that follows from the anchor result is routing infrastructure, not better interfaces. If the combination loses when the machine is the better component, then the first-order engineering problem is deciding which instances go to which party before the fact, and exposing that decision as an auditable artifact. Speculative Almost no deployed system has a routing layer of this kind; most have a recommendation layer and an override button, which is the architecture the meta-analysis measured and found wanting.

Frontier Second, instrumentation of the interaction rather than the output. Ju and Aral’s analysis of more than 550,000 messages is the first quantitative account of how the human’s role changes under augmentation — more delegation, far fewer direct edits — and it required building the collaboration platform to get the data. Established Systems that log only outcomes cannot detect the substitution of delegation for evaluation, which is the mechanism by which augmentation quietly becomes automation. Speculative Logging designed for second-derivative measurement — the rate at which a team’s own improvement rate changes — does not exist in any product this brief could identify, and Link 7 cannot be run without it.

Frontier Third, high-ceiling interfaces are an engineering choice that the market punishes, and the historical record is unusually clear about the cost of getting it wrong in either direction. NLS failed on modes, mnemonics, strict hierarchy and a chord keyset requiring users to learn a 5-bit binary code. Established Those are real design defects and no amount of philosophical vindication makes them not defects. Frontier But the professional tools that survive on Engelbart’s terms — modal text editors, spreadsheet formula languages, CAD and audio-workstation environments — demonstrate that a high training cost with a high ceiling holds durable markets where the users are professionals. Speculative The engineering lesson is that the co-evolution premium is a sectoral product, not a mass-market one, and building for both at once is what dissolved the original programme.

Frontier Fourth, calibrated confidence and honest failure signalling. Explanations raise acceptance without raising accuracy, so the useful signal is not why the machine said something but how likely it is to be wrong on this instance. Established The failure mode to engineer against is silent failure: a tool that is wrong without saying so, in a domain where the human’s residual skill has decayed, is the configuration in which every documented augmentation catastrophe has occurred.

8 · Adjacent technologies

Frontier The nearest neighbour is the hardware branch of the same ambition. Brain-computer work in Brain-Computer Interfaces and Neural Interfaces pursues bandwidth into and out of the nervous system, which is precisely the channel Engelbart proposed to route around rather than widen. Speculative If the augmentation ceiling is set by how much machine-generated signal a human can evaluate per unit time, then interface bandwidth is the binding physical variable and the software tradition has been optimising the wrong layer for sixty years. Handwave There is no measurement of that saturation point anywhere in the literature, which makes this the most testable-sounding of the exotic claims in this brief and one of the least tested.

Frontier The theory neighbours are distributed cognition and collective intelligence. Distributed Cognition supplies the unit of analysis the H-LAM/T construct needs — a human, artifacts and methodology treated as one computational system — and it arrived thirty years later without citing Engelbart. Frontier Collective Intelligence owns the homogenisation question, and the anchor meta-analysis came out of the same laboratory. Established The overlap of personnel is not a coincidence: the people who measure groups are the people who noticed that human-AI pairs are groups.

Frontier The substitution neighbours are where the waypoint argument gets decided. Artificial General Intelligence and Multi-Agent Intelligence Systems describe the trajectory that, on the waypoint reading, dissolves each augmentation configuration from below. Speculative If multi-agent architectures internalise the division of labour Licklider proposed between human and machine, they remove the structural reason for a human to be in the configuration at all, and augmentation survives only where human presence is required for reasons that are not accuracy reasons.

Established The oldest neighbour is human factors, and it has been answering these questions since 1983. Bainbridge’s “Ironies of Automation” established that automating most of a job while leaving the operator responsible for the remainder both erodes skill through disuse and imposes exhausting monitoring, so operators need more training rather than less. Frontier Every result in the current augmentation literature that surprises anyone was anticipated in that field, and the discipline’s vocabulary — automation bias, situation awareness, out-of-the-loop performance, alarm saturation — is better developed than anything the AI literature has built.

9 · Institutional requirements

Established The programme’s collapse was a funding-structure event and it is worth naming the mechanism precisely. The 1969 Mansfield Amendment ended military funding of non-military research; combined with the end of Vietnam and the termination of Apollo it drained ARPA and NASA support from the Augmentation Research Center through the early 1970s. Established No one adjudicated the science. Frontier The intellectual rejection came later and separately, from SRI management and from the researchers who left for Xerox PARC over accessibility, and the two causes are routinely merged into a single story about the vision failing.

Frontier The deeper institutional problem is that the co-evolution level of Engelbart’s model has no purchaser. Improving the process by which you improve your processes returns value on a horizon longer than an individual user’s patience, a manager’s tenure, or a vendor’s sales cycle, and vendors compete on time-to-first-value, which is exactly the metric co-evolution sacrifices. Speculative The strong version of that claim — that markets cannot price co-evolution — is probably false, because professional tool markets price it routinely. Frontier The sectoral version survives: mass markets cannot price it, and the augmentation tradition aimed at everyone.

Established The most consequential institutional requirement in this subject is a reporting standard, not a funding line. Requiring the AI-alone arm in any study claiming that AI assistance helps would have prevented a decade of comparisons against the wrong baseline, and it costs nothing beyond study design discipline. Frontier Pre-registration is arriving through trial registries in the education and economics literatures, which is the field attempting to get ahead of its own publication bias — and given that the positive augmentation finding is demonstrably bias-inflated at P = 0.002, that effort is not premature.

Speculative The Institute-relevant point is unusual and worth stating plainly. In most subjects the binding constraints are capability constraints. Here they are measurement-practice constraints: an unpublished taxonomy, a missing control arm, an unrun second-generation replication, an eighteen-month evaluation horizon nobody will fund. Frontier None of these needs new technology, a large budget, or a long research programme. Handwave A funder willing to spend on unglamorous replication rather than on capability could move this entire subject from opinion to evidence within about three years, and that opportunity appears to be sitting unclaimed.

10 · Ethical & societal considerations

Frontier The first ethical fact is that augmentation measurably erodes the skill it is supposed to augment, in at least one clinical setting. A 2025 multicentre observational study reported that adenoma detection in standard non-AI colonoscopy fell from 28.4% to 22.4% after routine exposure to an AI-assisted polyp-detection system. Frontier This brief flags the provenance honestly: the citation is verified against the bibliographic registry, the figure was read from a secondary source citing the paper, and the primary abstract was not obtained. Frontier The result generated four published correspondence responses and an authors’ reply within months, so it should be presented as contested rather than settled — but a contested first measurement of AI-induced deskilling is still the only measurement there is.

Established The second is automation bias, which has a settled taxonomy and hard numbers. Commission errors occur when an operator follows an automated directive without weighing contrary evidence, through overt redirection of attention away from the aid, diminished attention to it, or active discounting of information that contradicts it; omission errors occur when the operator fails to notice what the automation missed. Established In one cited breast-cancer study, cancers found in 46% of cases without an automated aid were found in only 21% of cases where the aid failed to identify them — a 25-point absolute drop caused by an aid that missed the finding. Frontier Training reduces commission errors but not omission errors, and training with deliberately injected failures works better than warning that failures are possible.

Speculative The third is that human-in-the-loop requirements may be buying legitimacy rather than accuracy, and it is better to say so. If combinations underperform the better component on decision tasks, then a regulatory apparatus mandating human oversight is procuring located responsibility, contestability and due process — which are genuine goods — while being defended on accuracy grounds it cannot support. Frontier Pretending otherwise sets such systems up to fail on their stated metric while succeeding on their actual one, and makes the arrangement impossible to improve because its purpose is misdescribed.

Frontier The fourth is distributional and cuts the other way. Every well-measured augmentation gain in this literature accrues to the weaker performer: the lowest-skilled support agents, the lowest-rated tutors, the least experienced clinicians. Speculative That makes augmentation an unusually progressive technology at the level of individuals and an unusually homogenising one at the level of populations, and those two facts are the same fact seen from different distances.

11 · Civilizational implications

Frontier If the waypoint reading is right, the augmentation era is a transitional regime and the institutions being built on it are mispriced. The relative-baseline moderator says the combination gains while the human is the better component and loses once the machine is, which means each augmentation configuration is dissolved from below as capability rises. Speculative On that reading “augmentation” names a moving boundary rather than an architecture, and the political economy currently organising itself around human-in-the-loop deployment is building on a receding shoreline.

Speculative If the equilibrium reading is right, the civilisational stake is the opposite one: the ceiling was never tested. Engelbart’s claim was that a group applying its improvement capability to its own improvement capability compounds, and that claim — a claim about second derivatives — has never been subjected to a controlled test in sixty years. Handwave If it holds even weakly, the returns to co-evolutionary tooling are compounding rather than additive, and the entire industry has spent six decades optimising a first derivative. Frontier There is no evidence for this and the absence of evidence is an absence of experiments rather than a record of failures.

Frontier The homogenisation risk is the one that scales badly and is measurable now. Augmenting an entire population with the same system raises individual performance and compresses output diversity; for problems that civilisations solve by parallel search — science, policy, culture — the second effect can in principle dominate. Speculative One controlled measurement of output self-similarity now exists. No population-scale measurement does. Handwave A civilisation could raise every individual’s floor and lower its own ceiling without any single actor being able to detect it, and nothing currently in place would notice.

12 · Timelines

These horizons track the empirical questions this brief identifies as decidable — whether complementarity is routable, whether it survives a capability generation, whether co-evolution has a premium — rather than the pace of capability, which is forecast elsewhere.

  • 10 yr: Frontier The task-class map is built and the sign of the combination effect is known for eight or more classes rather than two. Frontier At least one two-generation complementarity series is published; this brief expects the margin to shrink, and expects the first such publication to be contested on task-selection grounds. Speculative The AI-alone arm becomes a reporting norm in at least one professional literature, most plausibly clinical imaging, and the effect on published effect sizes is visible and downward. Speculative The self-report gap is replicated outside software, at which point organisational survey evidence for augmentation loses standing.
  • 25 yr: Speculative Either a routing policy that identifies the human’s instances in advance exists and augmentation becomes an engineering discipline with a design theory, or it does not and human-in-the-loop deployment is openly justified on governance rather than accuracy grounds. Speculative This brief judges the second more likely for decision tasks and leaves creation tasks genuinely open, the meta-analytic estimate there being indistinguishable from zero in either direction. Frontier Somebody has run Link 6 — the eighteen-month co-evolution cohort study — probably in a professional-training setting rather than a product one.
  • 50 yr: Speculative The IA/AI distinction has stopped carving anything and survives as a description of who holds the decision right in a given deployment. Speculative Augmentation persists robustly in exactly the places where human presence is required for non-accuracy reasons — accountability, preference elicitation, legitimacy — and has been dissolved everywhere else. Handwave If bootstrapping is real, the first organisations to have measured it have a compounding advantage by this point and are conspicuous; if it is not, Engelbart’s C-level is remembered as an elegant idea that no one could operationalise.
  • 100 / 250+ yr: Handwave Either the human channel-capacity ceiling is the permanent binding constraint on augmentation, in which case the only route past it is direct interface bandwidth and this subject merges into the neural-interface programme; or the ceiling is escaped by hierarchical delegation, in which case the human is supervising a supervision layer and the word “augmentation” has quietly come to mean ownership. Handwave The way to bet, on the evidence in this brief, is the second — and the interesting question then is not whether the tools amplify anyone but whether anything is left that the amplification is of.

13 · Technology tree & dependencies

  • Depends on This brief depends on Human-AI Integration for its central empirical result and for the entire teaming evidence base: the 106-study meta-analysis, the moderator structure, the automation-dependence record, the handoff failures, and the clinical reading studies are developed there. The dependency is unusually tight because a single number from that brief — the pooled effect of the combination against the better of its parts — determines whether the augmentation tradition is describing an architecture or a transient. If that estimate moves, or if the per-task-class extension finds classes with durable positive effects, this brief's central reading changes with it.
  • Enables What this subject enables is mostly a design discipline and a vocabulary rather than a capability. The H-LAM/T framing supplies the unit of analysis that any serious human-plus-tool measurement needs, and its insistence that language, methodology and training are co-equal with artifacts is the correction most current augmentation product work requires. The workback chain set out here — taxonomy, task-class map, pre-registered complementarity, ex-ante routing, generation survival, co-evolution premium, bootstrapping — is directly reusable by any field deploying decision support, and the routing requirement in particular is a concrete engineering brief that the education, clinical and software literatures all currently lack.
  • Adjacent Human-AI Integration, Collective Intelligence, Distributed Cognition, Human Cognitive Augmentation, Future Education Systems, Brain-Computer Interfaces, Neural Interfaces, Artificial General Intelligence, Multi-Agent Intelligence Systems. The adjacency that matters most is the first, and the one that matters most unexpectedly is distributed cognition, which reinvented Engelbart's unit of analysis thirty years later without citing him.

14 · Common misconceptions & speculative claims

Established “Two amateurs with three laptops beat grandmasters with supercomputers, which proves weak human plus machine plus better process beats strong machine.” This is the founding anecdote of the entire human-AI complementarity discourse, and it does not survive contact with the record. Established What is sourced: Advanced Chess was created by Kasparov in 1998; freestyle chess permits every form of consultation within the time limits; the PAL/CSS Freestyle Tournament ran from 2005, organised by Computer-Schach und Spiele on ChessBase’s Playchess server with prizes totalling 132,000 euros between 2005 and 2008; and an amateur expert in chess software won, to the surprise of many who were certain the grandmasters would prove superior. Frontier What is not sourced is almost everything the story is told for. Established A bibliographic search across the scholarly record returned zero empirical studies of centaur or freestyle chess performance; every hit on “centaur” was metaphorical, in open textbooks and stock-trading papers. Established The winners’ names differ between retellings: the encyclopedic account gives Steven Cramton and Stephen Zackery, the widely circulated Kasparov version Steven Cramton and Zackary Stephen. Frontier The ratings universally quoted for the pair appear in no source this brief could read and are not printed here. Established The dispute over whether the advantage still exists is on the record with dates: Tyler Cowen wrote in 2013 that engine advances had made any major centaur advantage hard to see and unlikely to last; Kasparov in 2017 and James Bridle in 2018 assert the contrary, in trade books, without data. Speculative The most-cited exemplar of human-AI complementarity rests on one tournament series that ended in 2008, has produced no peer-reviewed analysis in nearly two decades, cannot agree on who won it, and carries unverifiable ratings. Tell it as a story and say that is what it is.

Established “The chess case at least shows the human still adds something.” The context says otherwise. Komodo was rated 3361 by the SSDF in 2016; FIDE no longer accepts human-computer results in its rating lists; grandmaster Andrew Soltis’ 2016 verdict was that the computers are just much too good; and players now treat engines as analysis tools rather than opponents. Frontier One measured centaur-versus-engine result after 2020 would settle it and, as far as this brief can determine, none exists.

Established “The meta-analysis showed human-AI teams help on creative tasks and hurt on decision tasks.” Half right. Decision tasks: g = -0.27, P = 0.002, 95% CI -0.44 to -0.10. Creation tasks: g = +0.19, P = 0.180, 95% CI -0.09 to 0.48 — the interval spans zero. Frontier The abstract’s phrase about significantly greater gains in tasks involving content creation is a statement about the contrast between task types, and is routinely read as a statement about the creation effect itself. This is the commonest misreading of the paper.

Established “That meta-analysis is negative, but this literature is biased toward positive findings anyway, so the true picture is better.” The bias runs the other way, and this is the most under-quoted result in the paper. For the negative primary comparison, Egger’s regression found no evidence of publication bias (beta = -0.67, P = 0.438) and neither did the rank correlation test (tau = 0.05, P = 0.121). For the positive comparison against humans alone, Egger’s regression does indicate bias (beta = 1.96, P = 0.002). Frontier The pessimistic headline is bias-clean and the optimistic g = 0.64 is inflated, so any argument discounting the negative finding on publication-bias grounds has the paper backwards.

Established “Intelligence amplification is Engelbart’s term.” It is Ashby’s, from 1956, and it is a claim about selection ratios rather than about computers. Established “Engelbart invented the mouse and personal computing.” He demonstrated the mouse as one of ten items in a ninety-minute demo whose thesis was distributed collaborative knowledge work, conducted over two live microwave video links to a laboratory thirty miles away. Frontier Artifacts are one of his four classes of augmentation means and, on the tradition’s own account, not the primary one — though this brief could not obtain the 1962 report and therefore cannot confirm Engelbart’s own ordering. Speculative If the report weights artifacts first, the standard critique of how the demo was received collapses.

Established “The 1968 demo showed the future and the industry built it.” The industry built the artifacts and dropped language, methodology and training. Established His system was explicitly not built for ease of use — WYSIAYG, What You See Is All You Get, was his own coinage for the position — and his colleagues left for Xerox PARC precisely to abandon that stance. Frontier What was built was the negation of the design philosophy, using the widgets. Speculative Whether that cost anything measurable is unknown, because the experiment that would show it has never been run.

Frontier “Licklider predicted human-computer partnership as the enduring condition.” The secondary record says the opposite: the paper positions symbiosis as an interim arrangement pending machines outdoing the human brain in most of the functions it describes, and enthusiast citations routinely truncate this. Handwave This brief does not publish that as a correction. The exact wording and the interval Licklider gave could not be verified, the primary text was not obtainable, and a claim that a founding paper denies the result it is cited for is precisely the kind of claim that must be read before it is made. Frontier It is recorded here as the highest-value open verification task in this subject: pages 4 to 11 of HFE-1, March 1960. If it holds, it belongs in the opening paragraph of this brief rather than in section 14.

Established “The augmentation vision failed for lack of hardware.” Alan Kay, who had the hardware, has held for decades that it is already here and that the missing pieces are key software and educational curricula — his verdict on Microsoft’s 2001 tablet was that it was the first Dynabook-like computer good enough to criticize. Established The documented failure account for NLS is learning curve, modes, mnemonics, chord keyset and a refusal to optimise for accessibility; hardware appears nowhere in it. Established “ARC was defunded because the vision was rejected.” The documented causes are the 1969 Mansfield Amendment, the end of Vietnam and the end of Apollo. It was macro-budgetary; the intellectual rejection came later and from a different direction.

Established “IA is the humane alternative to AI.” The split hardened as a competition for the same ARPA money, and Engelbart’s laboratory was handed to an artificial-intelligence researcher who negotiated its disposal. Speculative The distinction as the tradition draws it — that IA needs technology merely as extra support for an autonomous intelligence that has already proven to function — is a claim about which component is trusted, which is a deployment choice rather than an architectural one, and the premise it rests on is exactly what the decision-task evidence now contests.

Handwave “Augmentation is obviously better than automation.” The largest meta-analysis of the question finds combinations averaging below the better component at g = -0.23, on the comparison that is clean of publication bias. Speculative The harshest available reading goes further: that “augmentation” is unfalsifiable, because any deployed system containing a human can be described as augmentation, and the word therefore functions to make automation politically legible rather than to describe an architecture. Frontier The anchor corpus is drawn entirely from studies self-described as human-AI collaboration and averages negative against the better part — a literature calling substitution-with-friction augmentation. Frontier The reply is that this is a definitional failure rather than an empirical one, and the fix is Link 1 of the workback plan in section 6: an operational definition that excludes some substantial class of human-in-the-loop deployments. It is cheap, it is undone, and until it exists this subject cannot say precisely what it is about.