1 · Concept overview
The framing under test contains a hidden conjunction. “AI raises productivity” and “measurably” are separate claims with separate evidence bases, and as of 2026 they point in different directions. The firm-level and task-level evidence for large effects is strong, replicated and growing: a peer-reviewed staggered rollout across five thousand customer-support agents, a randomised trial on general knowledge work, a clean single-firm study of taxi dispatch, and a hundred thousand developers across three generations of coding tools. The aggregate evidence is close to absent. And the people who have looked hardest at whether the aggregate statistics could detect an effect disagree about whether they could — with one economist appearing on both sides of that disagreement, on two different questions that are routinely merged.
The arithmetic that dissolves most of the apparent paradox is available and rarely stated. Generative AI adoption in the United States is historically fast — faster than the personal computer or the internet at the same point — and the share of work hours that are actually AI-assisted is in the low single digits. The product of a large per-task effect and a small assisted share is a small aggregate effect. That is not a mystery requiring lags or mismeasurement to explain; it is multiplication.
This brief owns the productivity term of the task model: output per hour, the firm-level effect sizes, the aggregation problem between task gains and shipped output, intangible capital, and whether national accounts can see any of it. The displacement term — employment, wages, the labour share and who bears the adjustment — belongs to Future Labour Markets. Where a source here also carries an employment number, the number is named and handed over rather than interpreted.
2 · Current scientific position
Established Start with the null that has the best claim on the reader's attention, because the interest runs against it. Yotzov and co-authors surveyed nearly 6,000 senior executives across the United States, the United Kingdom, Germany and Australia. 69% of their firms actively use AI. Asked about the preceding three years, nine in ten reported no impact on employment or productivity. Asked about the next three, the same executives forecast productivity +1.4%, output +0.8% and employment −0.7%. Their own personal use of AI averages 1.5 hours a week. Established Senior executives have every incentive — to boards, to investors, to their own narrative — to report that an AI programme is working, and nine in ten said it had not yet done anything. Frontier The retrospective and the prospective answers come from the same people in the same instrument, which makes the gap between them the cleanest available measurement of how much of the AI productivity story is anticipation rather than observation.
Established That null survives the diffusion excuse, which is what makes it different from the usual “it is early” exchange. These are not non-adopters. If the answer is still lags, the lag is operating after adoption — a different and much less comfortable claim than “the technology has not arrived yet”. Frontier A second executive survey disagrees, and the disagreement is itself the state of the evidence: Baslandze and co-authors, on roughly 750 corporate executives, report positive productivity gains varying by sector and concentrated in high-skill services and finance, attributed to total factor productivity rather than capital deepening, with little evidence of near-term aggregate employment decline. Established That paper also reports the finding that reconciles the two: executives perceive productivity gains as larger than measurable gains, which its authors attribute to delayed revenue realisation. Two executive surveys, different samples, opposite signs on the retrospective question.
Established The firm-level and task-level effects are real, large, and narrower than they read. Brynjolfsson, Li and Raymond, on a staggered rollout of a generative AI assistant across 5,179 customer support agents, published in the Quarterly Journal of Economics: +14% productivity on average, +34% for novice and low-skilled workers, and minimal effect on experienced high-skilled workers, with improved customer sentiment, improved retention and suggestive evidence of worker learning. Established Cruces and co-authors randomise general knowledge work across 1,174 adults: without AI, higher-education participants outperform lower-education participants by 0.548 standard deviations; with AI access the gap falls to 0.139, closing about three-quarters of it; removing AI leaves lower-education participants retaining part of the gain but a sizeable gap re-emerging. Frontier Japanese taxi drivers given AI demand prediction gain through shorter cruising time, and the gains accrue only to low-skilled drivers, narrowing the skill gap by 14 percentage points. Four independent settings, one pattern: the effect is concentrated in the lower half of the skill distribution.
Frontier The most important number in this brief is an attenuation, not an effect size. Demirer, Musolff and Yang track over 100,000 GitHub developers across three generations of AI coding tools and find cumulative effects on commits of +40% for autocomplete, +140% for interactive agents and +180% for autonomous agents. Then they follow the gain downstream: that 180% on commits falls to 50% on number of projects and 30% on actual releases. They attribute the attenuation to a weak-link mechanism — human bottlenecks moving downstream of the automated step — and estimate an elasticity of substitution between AI and human effort of 0.25, which is strong complementarity rather than substitution. Frontier It is observational rather than randomised and it is a 2026 result on one platform. Established But it is the only fetched study that follows a measured AI gain all the way to a shipped product, and five-sixths of the magnitude is gone by the time it arrives. Every other headline in this literature is measured on an intermediate output: commits, tickets resolved, words written.
Established Firm adoption is far lower than consumer adoption, and firm adoption is what the aggregate statistics respond to. McElheran and co-authors, using the 2018 Annual Business Survey across roughly 850,000 US firms, find fewer than 6% of firms used any of five AI-related technologies, rising to just over 18% on an employment-weighted basis, with most very large firms reporting some use and adoption concentrated in a handful of superstar cities. Frontier That is a 2018 snapshot and it predates generative AI, so it establishes a shape rather than a current level: AI use is employment-weighted, not firm-weighted. Established On the current wave, nationally representative surveys find nearly 40% of the US population aged 18–64 using generative AI by late 2024, 23% of employed respondents using it for work at least weekly and 9% daily, with work adoption as fast as the personal computer and overall adoption faster than either the PC or the internet. Frontier The same work estimates only 1–5% of all work hours are currently AI-assisted, with reported time savings of about 1.4% of total work hours — self-reported, which is a weak instrument. That pair of numbers is the whole subject in miniature: adoption historically fast, assisted share small.
Frontier The most disciplined bottom-up estimate is modest and its method is checkable. Acemoglu runs Hulten's theorem over the share of tasks AI can plausibly perform and the cost saving on each, and gets no more than a 0.66% increase in total factor productivity over ten years, and less than 0.53% once harder-to-learn tasks are handled conservatively. He calls it “nontrivial but modest”; it is published in Economic Policy. Frontier The arithmetic is transparent; the inputs — the share of exposed tasks and the average cost saving — are judgement calls that drive the answer. Established The structure of the disagreement with the bullish projections matters more than either point estimate. The bottom-up method enumerates tasks, applies measured per-task effects and aggregates under a theorem saying the aggregate gain is the cost-share-weighted average of the micro gains. The top-down method assumes a share of labour is automatable, assumes it is automated, and prices the labour. Handwave The two cannot be reconciled by argument, because they disagree about whether the relevant input is measured cost savings on tasks AI actually performs or the wage bill of tasks AI could in principle perform.
Frontier And one dissent comes from inside the framework, with the sign nobody expects. Acemoglu and Restrepo's rent-dissipation work estimates that automation-driven rent dissipation is inefficient and has offset 60–90% of the productivity gains from automation since 1980, and “could even negate” them. Frontier If that carries over to AI, the modest bottom-up figure is an upper bound on realised productivity rather than a central case. Established The same paper's inequality result belongs to Future Labour Markets and is not restated here. Speculative The claim is recent and the field has not absorbed it; it is carried because a mechanism that turns a positive into a near-zero deserves stating even before it is digested.
3 · Frontier questions
Frontier Is it lags? The 2017 statement of the paradox set out four candidate explanations — false hopes, mismeasurement, redistribution and implementation lags — and argued lags were likely the biggest contributor: general purpose technologies need complementary innovations and organisational change before they show up. Frontier Nine years on the paradox is intact in a stronger form, because the diffusion excuse is weakening: adoption is high, and the null is being reported by adopters. Frontier What would settle it is aggregate productivity acceleration by roughly 2030, given that mass generative-AI adoption dates from 2023 and the personal computer took roughly fifteen years from mass adoption to the late-1990s acceleration.
Frontier Is modest simply the right answer? The bottom-up accounting says at most two-thirds of a percentage point of total factor productivity over a decade, and it is the estimate most consistent with the fetched aggregate evidence — and with the measured assisted-hours share. Frontier What would settle it is better measurement of the share of tasks actually performed by AI, which currently rests on self-report.
Frontier Is it mismeasurement? The J-curve argument says complementary intangible investment is poorly captured in the national accounts even when it creates valuable assets, so measured productivity is understated during the build phase and overstated during the harvest. Frontier Against it stand four arithmetic challenges, none rebutted in the fetched literature. This hypothesis is live and trending weaker, and the reason it is trending weaker is section 4.
Frontier Is it the weak link? Task gains are real and large but do not aggregate, because the binding constraint moved downstream to humans. The evidence is the commits-to-projects-to-releases attenuation and an elasticity of substitution of 0.25. Frontier This brief's assessment is that it is the most under-rated hypothesis in the subject, because it explains both the large randomised effects and the aggregate null without invoking lags or mismeasurement. What would settle it is replication of the intermediate-versus-final-output decomposition outside software, which nobody has done.
Frontier Is it levelling rather than raising? Four settings replicate a pattern in which AI compresses the skill distribution instead of shifting the mean: +34% for novices against minimal for experts, gains only for low-skilled drivers, a 0.548 standard-deviation education gap closing to 0.139. Speculative The aggregate implication is speculative and uncomfortable: if the gain accrues to below-median performers, the economy-wide effect depends on the share of output those performers produce, which in most industries is small. An intervention that raises the floor and leaves the ceiling can be transformative for individuals and nearly invisible in a national account.
Frontier Is it false hopes — a good interface to existing information rather than a productivity technology? The evidence is the executive retrospective null at 69% adoption, one and a half hours a week of personal executive use, and assisted hours in single digits. Speculative A sharper minority version holds that expert task-level effects may be negative: that experienced practitioners are slower with AI assistance while believing themselves faster. Handwave The study most often cited for that could not be verified in this pass, so it is stated as a hypothesis with a named consequence — if it generalises, every self-reported time saving in this literature is biased upward, including the 1.4% figure and every executive survey.
Frontier Is the story the cost structure rather than the productivity? Foundation models have very large fixed costs in training, compute and talent and non-trivial marginal costs in inference, producing concentration rather than diffusion. Speculative If the rents accrue to a few model providers, measured aggregate productivity can stay flat while enormous value is created and captured. Frontier That is a market-structure claim before it is a productivity claim, and it is developed on the market-structure side in Digital Economies; the two briefs hold the same hypothesis from opposite ends.
Speculative And the fringe position, stated because it is coherent: measured productivity is the wrong object entirely. If AI's output is predominantly free or near-free goods, gross domestic product is definitionally blind to it and the right measure is a welfare-based one. Frontier The counter is the magnitude challenge: existing estimates of consumer surplus from digital technologies are far too small to close the gap. The fringe reading is not obviously wrong; it is arithmetically insufficient, which is a different objection and a stronger one.
4 · Technological bottlenecks
Established The intangible-capital measurement gap is the binding constraint, and it is institutional rather than technical. General purpose technologies require complementary investment — process redesign, business-model change, human capital — that creates valuable assets and is not capitalised in the national accounts. During the build phase costs are recorded and benefits are not, so measured productivity is understated; during the harvest phase it is overstated. Frontier The J-curve estimates put adjusted total factor productivity 11.3% higher than official measures at the end of 2004 and 15.9% higher by the end of 2017, with the largest effects for software. Frontier Those estimates rest on an assumed relationship between firm market value and an unobserved intangible stock, which is the step a sceptic should press on. The gap itself is not in dispute; its size is.
Frontier The second measurement gap is that nobody counts AI-assisted hours. Statistical agencies count adoption — whether a firm reports using a technology — because that is what business surveys were built to ask. The variable that determines the aggregate effect is the share of work hours in which the technology is actually used, and the only estimates of it are survey self-reports in the 1–5% range. Established Adoption and intensity have diverged before, and the divergence is exactly what a firm-weighted-versus-employment-weighted gap looks like: fewer than 6% of firms, just over 18% of employment. Frontier Until an official series measures intensity, “adoption is high” and “usage is deep” will keep being said as if they were the same sentence.
Established The third bottleneck is that almost every measured outcome is an intermediate one. Commits, tickets resolved, words produced, cruising time saved. A firm's productivity is output per hour of things it actually sells. Frontier The single study that bridges the gap loses five-sixths of its effect in the crossing, and the mechanism it proposes — the constraint relocating to an unautomated human step downstream — predicts the same loss anywhere the automated step was not the binding one. Speculative If that is general, then the published effect sizes are measurements of where the slack was, not of how much output rose.
Frontier Fourth: nulls are under-published, and the one large null in the corpus is a survey rather than an experiment. The randomised evidence is uniformly positive and uniformly conducted in settings chosen because a task looked amenable. Speculative A demonstration that the firm-level effects are selection on task type — that randomly selected tasks show much smaller effects than the tasks chosen for study — would be the single most informative experiment available, and nobody has run it. Established Publication incentives here run in the direction of positive findings on both sides of the market: researchers want effects and vendors want evidence.
Established And self-reported time savings are not a productivity measurement. They measure a belief about time saved. No fetched source validates them against output, and if the expert-slowdown hypothesis is right they are biased upward exactly where AI use is most sophisticated. Frontier Every headline estimate that rests on them — including the 1.4% of work hours figure — should be read as an upper bound on a belief, not as a lower bound on an effect.
5 · Research dependencies
Established Nothing on this map produces a result this brief waits on, and no typed depends-on edge is claimed. What it waits on is two measurement systems and one missing study, all three recorded as typed requirements in section 13. Two of them a statistical agency could choose to build. The third is a result nobody is currently producing.
Established The first dependency is national accounts that capitalise intangible investment. This is not a plea for better statistics in general; it is a specific accounting treatment with a known sign and a known timing profile. Until organisational capital, process redesign and firm-specific human capital are capitalised rather than expensed, measured productivity will be understated during exactly the phase this technology is in, and no amount of better firm-level evidence will resolve an argument about the aggregate.
Established The second is a statistical series for AI-assisted work hours. The quantity that determines whether task-level gains become aggregate gains is intensity, and no official instrument measures it. The existing estimates are self-reported, are the weakest link in the chain, and are the ones every forecast is implicitly scaled against.
Frontier The third is a firm-level total-factor-productivity effect identified off something exogenous — a supply shock in compute access, a licensing-cost discontinuity, a staggered vendor rollout across otherwise comparable firms. Established The customer-support study is the template at the level of a single occupation inside a single firm; nothing at that quality of identification exists at firm scope. Speculative Until it does, the field is comparing randomised task-level effects against observational aggregates and calling the difference a paradox. This is recorded as a scientific requirement rather than an institutional one because no agency can commission a natural experiment into existence.
Established What this brief explicitly does not wait on is the employment question. Who gained work, who lost it, and what happened to the labour share are adjudicated in Future Labour Markets. The two briefs split a shared literature at the model's two terms and neither re-derives the other's half.
6 · Required experiments
Established The highest-value experiment is the cheapest and has been specified for two years: replicate the intermediate-to-final-output decomposition outside software. Take any setting with a measured AI gain on an intermediate output and follow it to a shipped, sold or delivered unit. Customer support has tickets and renewals; radiology has reads and diagnoses; legal drafting has documents and matters closed. Frontier If the five-sixths attenuation recurs, the aggregate null is explained without lags or mismeasurement and the field can stop arguing about both.
Frontier Second: randomise the task, not the tool. Every published effect size comes from a setting selected because the task looked amenable to AI. Drawing tasks at random from a firm's actual work and randomising assistance across them would measure the quantity that matters for aggregation — the average effect over the work a firm does, rather than the effect on the work a researcher chose. Speculative Nobody has done this, and the expected result is a substantially smaller number.
Established Third: re-survey the same executives. The nine-in-ten retrospective null is paired with a +1.4% three-year forecast from the same instrument. Re-running it in 2028 converts a snapshot into a scored forecast, at trivial marginal cost, and it is the single cleanest test available of whether the anticipation is being realised. Frontier A forecast made by named respondents with a stated horizon is a rare and perishable asset in this literature, and letting the horizon pass without re-measuring would waste it.
Frontier Fourth: find exogenous variation in AI access at firm scope. Compute allocation shocks, enterprise licensing thresholds, regional model-availability differences and staggered vendor rollouts are all candidate instruments, and at least the last of them is generated routinely by vendors without anyone recording it as an experiment. Speculative The obstacle is not method; it is that the variation is held by firms with no incentive to publish it.
Established Fifth: measure assisted hours in an official instrument. Adding two questions to an existing business survey — what share of employee hours involve AI assistance, and on which tasks — would replace the weakest number in the subject with a measured one. Frontier It is a questionnaire change, not a research programme, and its absence is the reason the diffusion argument cannot currently be adjudicated.
Speculative And sixth, the experiment that would test the uncomfortable hypothesis: run a properly powered trial on experienced practitioners with objective output measurement and simultaneous self-report. If experts are slower while believing themselves faster, the gap between the two measurements is the correction factor for every self-reported estimate in this brief. Handwave The result that would motivate this could not be verified here, which is precisely why the experiment is worth naming rather than citing.
7 · Engineering requirements
Established The engineering problem revealed by the attenuation result is that automating a step relocates the constraint rather than removing it. Writing code got much faster; shipping code did not, because review, testing, integration, release management and the humans doing them did not. Frontier The measured elasticity of substitution of 0.25 says AI and human effort are strong complements in this setting, which means the return to the next unit of AI depends on whether the complementary human capacity exists — and that is a staffing and process fact, not a model capability fact.
Established The complementary investment is the expensive part, and every large measured effect comes from a setting where it had already been made. The customer-support deployment sits on a firm with a ticketing system, a quality-scoring apparatus, a training pipeline and a measurable output. Frontier A firm without those cannot buy the +14%, because the instrument that produced it is the surrounding system rather than the assistant. Speculative This is the strongest version of the lags argument and the one least often stated: the lag is not diffusion of the technology, it is construction of the measurement and process apparatus that lets a firm both use it and see the result.
Frontier Inference has a real marginal cost, which changes the deployment calculus. Foundation models carry very large fixed costs in training, compute and talent and non-trivial per-query costs. A tool that costs money every time it is used is deployed where the value per query is highest, which is not where the hours are. Speculative If the assisted-hours share stays in single digits while adoption approaches saturation, a per-query cost structure is a sufficient explanation and requires no appeal to organisational inertia.
Established And the skill-composition result has a direct engineering reading. The gain is a floor-raising one: it converts a novice into a median performer and does little for an expert. Frontier A firm's realisable gain therefore depends on the shape of its internal performance distribution, which nobody measures and which varies enormously across industries. Deployment planning built on an average effect size is planning against a statistic that does not describe any particular firm.
8 · Adjacent technologies
Within this map: Future Labour Markets, the other half of the same task model and the owner of every employment and wage number named here; Digital Economies, which develops the AI cost-structure hypothesis from the market-structure side; Innovation Ecosystems, where the public-investment case for AI research is evaluated against its evidence; Artificial General Intelligence and Multi-Agent Intelligence Systems, which describe the capability trajectory this brief measures the economic consequences of; Artificial Scientists, the one deployment where an intermediate-output measure and a final-output measure might genuinely coincide; Ultra Efficient Computing Energy Systems, which carries the physical cost of inference; and Future Public Administration, whose finding that evaluation is discretionary explains why no statistical agency has been required to measure any of this.
Outside it: growth accounting and the measurement of total factor productivity; the economics of general purpose technologies; national accounting practice, which is where the intangible-capital question is actually decided; and the empirical literature on organisational capital, which supplies the complementary-investment argument that both sides of the paradox rely on.
9 · Institutional requirements
Established The institutional constraint here is a national accounting convention, and it is unusually specific. Intangible investment that creates durable assets — organisational capital, process redesign, firm-specific training — is largely expensed rather than capitalised. That treatment has a known sign at each phase of a general purpose technology's diffusion, and it is being applied during the phase where it understates. Frontier Changing it is a decision by statistical authorities with published methodological standards, not a discovery, and the estimates of what it would change range from substantial to nearly nothing depending on whose method is used.
Established The second institutional gap is that no official instrument measures usage intensity. Business surveys ask whether a technology is used because that is what they were designed to ask about machine tools. The variable that governs aggregation is the share of hours, and it is currently supplied by academic self-report surveys. Frontier A statistics office finding its own productivity series unable to detect a technology at 69% firm adoption would be exactly the interest-against-finding evidence this subject most needs, and this pass located none.
Established Interested parties are heavily represented on one side of this argument and the marking matters. The bullish multi-point growth projections in circulation come from consultancies selling AI transformation advisory and from institutions with equity exposure to AI capital expenditure; the interest runs with the forecast. Established None of them could be fetched in this pass, so none is characterised here beyond its method — and its method, assuming a share of labour is automatable and pricing it, is the top-down approach whose disagreement with the bottom-up approach is described in section 2. Frontier On the other side, vendor-authored studies of coding assistants exist and are cited widely; where they are eventually used they must be marked as vendor research.
Frontier And the evidence base has a structural asymmetry no institution is fixing. The randomised evidence comes from firms that agreed to be studied, on tasks selected for amenability, with the complementary systems already built. The aggregate evidence comes from everyone. Speculative A registry of AI deployment trials, including abandoned ones, would do for this subject what trial registration did for clinical research, and nothing in the fetched corpus suggests anyone is building one.
10 · Ethical & societal considerations
Established The distributional shape of the measured gains is the opposite of the usual worry, and it is replicated. In four independent settings the benefit accrues to novices, low-skilled workers and the less educated, and barely at all to experts: +34% against minimal in the same job, gains confined to low-skilled drivers, an education gap closing by three-quarters. Frontier If that holds at scale it is a compression of the earnings distribution rather than a widening — which is not what most public argument about AI assumes, and which sits awkwardly beside the displacement findings in Future Labour Markets. Speculative The two are not contradictory: a technology can compress the distribution among those who keep the work while removing the work from others, and no fetched source measures both margins together.
Frontier The retention result deserves stating because it is the least-cited finding in the most-cited study. The customer-support deployment improved employee retention and customer sentiment alongside output. Speculative Whether that generalises or reflects one firm's implementation is untested, and it is exactly the kind of secondary outcome that stops being measured once a technology is routine.
Established The honest ethical problem is that the argument is currently unfalsifiable in practice, and that has consequences for people. Investment decisions, staffing decisions and public policy are being made against forecasts whose authors have no exposure to being wrong, using a method that cannot be reconciled with the method that produces modest numbers. Handwave “It is early” is compatible with every observation available today, which makes it a handwave rather than a hypothesis — and it is currently the field's most common response to the null.
Frontier And if the rent-dissipation mechanism carries over, the ethical accounting inverts. A technology that offsets most of its own productivity gains through inefficient reallocation is not producing a surplus to be distributed; it is producing a transfer with a small residual. Speculative That claim is recent, unabsorbed and contested. It is carried because a policy debate conducted entirely about how to share the gains should know that a serious estimate says most of the gains are dissipated in the taking.
11 · Civilizational implications
Established The terminal position is a declared tie between two well-evidenced halves, and the tie is the finding. AI raises productivity at the level of a task, a worker and a firm: this is measured, peer-reviewed and replicated across four independent settings. AI has not raised measured productivity at any level the national accounts can see: no fetched source reports an aggregate effect, and the largest survey of adopters reports nine in ten seeing nothing over three years. Established Both are true. Picking one requires ignoring the other, and most public argument does exactly that by choosing a level of analysis and calling it the answer.
Established The arithmetic that connects them is not in dispute and is rarely stated. A large per-task effect multiplied by a small assisted-hours share is a small aggregate effect. At 1–5% of hours, even a 30% task gain buys well under a percentage point of aggregate output. Frontier The interesting question is therefore not whether the effect is real but what governs the assisted share — per-query cost, task selection, complementary process capacity, or the human step downstream that the gain runs into. Each of those is a different forecast, and the public argument has mostly not noticed that it is choosing between them.
Frontier The civilisational risk is not that the gains fail to arrive; it is that the instruments cannot adjudicate the question for another decade. A society can invest a substantial share of its capital formation in a technology whose aggregate return it cannot measure, on the strength of firm-level trials in settings chosen for amenability, and discover the answer only when the diffusion window closes. Speculative The window has a date attached: continued flat aggregate productivity through 2028–2030 with assisted hours above half would retire the lag explanation, and the assisted-hours share staying in single digits while headline adoption rises would mean adoption is shallow rather than early. Established Those are stated in advance and they are cheap to check, which is the most useful thing this brief can leave behind.
Speculative And there is a version of this in which measured productivity is simply the wrong instrument for the century. If a growing share of what a technology produces is free at the margin, welfare and output diverge permanently, and a civilisation optimising the measured series optimises the wrong one. Frontier The available counter is arithmetic rather than conceptual — existing consumer-surplus estimates are far too small to close the gap — and it is a strong counter. Handwave But “the measure is wrong” and “the measure is right and the effect is small” make identical predictions about the series everyone watches, and distinguishing them requires building the other measure first.
12 · Timelines
These horizons track diffusion, statistical practice and one dated forecast rather than model capability:
- 10 yr: Frontier The decisive window is 2028–2030, and it is decisive because it was named in advance: mass generative-AI adoption dates from 2023 and the personal computer took roughly fifteen years from mass adoption to acceleration. Established The executive forecast of +1.4% productivity has a three-year horizon and named respondents, so it can be scored rather than argued about. Frontier Expect the weak-link hypothesis to be tested outside software inside this window, because the test is cheap and the result would settle the largest open question. Speculative Expect no official series for AI-assisted hours unless a statistical agency is instructed to build one.
- 25 yr: Speculative If the effect is real and modest, this is the horizon over which a 0.5–0.7% total-factor-productivity gain becomes visible in a series with a standard error larger than itself — which is to say it may never be separable from everything else happening at the same time. Frontier The intangible-capital treatment is likely to change at some point in this window on its own methodological schedule, and when it does, part of the apparent acceleration will be an accounting revision rather than an economic event. Handwave Anyone who cannot say in advance how much of a future acceleration they would attribute to the revision is not making a forecast.
- 50 yr: Speculative At this horizon the object of interest is whether the cost structure diffused or concentrated. If inference stays expensive and scale economies dominate, the technology looks like a capital-intensive utility and its gains show up as rents to a few providers rather than as economy-wide productivity. Frontier That is the same hypothesis Digital Economies holds from the market-structure side. Handwave Which way it resolves is a question about entry conditions and semiconductor economics, and extrapolating either is not a measurement.
- 100 / 250+ yr: Handwave Beyond useful forecasting. The only relevant base rates are the two previous general purpose technologies whose productivity effects took decades to appear and were then argued about for decades more. Handwave Two episodes is not a distribution, and the honest position is that a century-scale claim about this technology is a genre rather than an estimate.
13 · Technology tree & dependencies
- Depends on Nothing on this map. No result another brief produces is a precondition for this one: the firm-level evidence exists and is good, and what is missing is measurement infrastructure and one identification design. No typed depends-on edge is claimed.
- Requires (not on this map) Two accounting decisions and one missing study. First, national accounts that capitalise intangible investment: organisational capital, process redesign and firm-specific human capital create durable assets and are largely expensed rather than capitalised, which understates measured productivity during exactly the build phase this technology is in and overstates it later — the J-curve estimates put adjusted total factor productivity 11.3% above official measures at end-2004 and 15.9% above by end-2017, on an assumption about the relationship between firm market value and an unobserved intangible stock that a sceptic should press. Second, an official series for AI-assisted work hours: business surveys measure adoption because they were designed to ask about machine tools, while the variable that governs whether task gains become aggregate gains is intensity — currently supplied by academic self-report in the 1 to 5% range, the weakest number in the subject and the one every forecast is implicitly scaled against. Both are decisions a statistical authority could take. The third is not: an estimate of AI's effect on firm-level total factor productivity identified off exogenous variation in access — a compute supply shock, a licensing-cost discontinuity, a staggered vendor rollout across otherwise comparable firms. The customer-support study is the template at the level of one occupation inside one firm, and nothing of that identification quality exists at firm scope. Until it does, the field compares randomised task-level effects against observational aggregates and calls the difference a paradox. No agency can commission a natural experiment into existence, which is why this is recorded as a scientific requirement rather than an institutional one.
- Enables In principle every economic assessment on this map that assumes a productivity dividend from automation — the labour, investment and public-finance questions in this category, and the deployment cases elsewhere — depends on whether that dividend is real and measurable. No typed enabling edge is claimed, because the enabling relationship runs through an aggregate effect that no fetched source has observed. A brief cannot enable another on the strength of a result it reports as absent.
- Adjacent Growth accounting and the measurement of total factor productivity; the economics of general purpose technologies; national accounting practice, where the intangible-capital question is actually decided; the empirical literature on organisational capital; and within this map Future Labour Markets, Digital Economies and Artificial General Intelligence.
14 · Common misconceptions & speculative claims
Handwave “AI will add several percentage points to annual growth.” Every projection of that magnitude in circulation is produced top-down: assume a share of labour is automatable, assume it is automated, price the labour. Established The bottom-up alternative enumerates tasks, applies measured per-task cost savings and aggregates under a theorem, and returns at most 0.66% of total factor productivity over ten years. Established The two methods cannot be reconciled by argument, because they disagree about whether the relevant input is measured savings on tasks AI performs or the wage bill of tasks AI could perform. Established And the provenance matters: the bullish forecasts come from firms selling AI transformation advisory or holding equity exposure to AI capital expenditure. The interest runs with the forecast, and none of them could be fetched in this pass, so none is characterised here beyond its method.
Established “Adoption is high, therefore the effect is arriving.” 69% firm adoption coexists with nine in ten adopters reporting no productivity or employment impact over three years. Established Consumer adoption is faster than the PC or the internet, and the share of work hours that are AI-assisted is 1–5%. Frontier Adoption and intensity are different variables and only one of them multiplies into output. The gap between them is the single most common error in this subject and it is arithmetic, not judgement.
Established “+14% in customer support means +14% everywhere.” The same paper reports minimal effect on experienced high-skilled workers in the same job, and +34% for novices. Established Across four settings the gain is concentrated in the lower half of the skill distribution, so the transferable quantity is not an effect size but a rule: the benefit tracks the share of output produced by below-median performers. Speculative In most industries that share is small, which is a reason to expect the aggregate effect to be smaller than the trial effects rather than merely delayed.
Frontier “The gains are there and the statistics cannot see them.” This is a real hypothesis with a real quantitative estimate behind it, and it faces four unrebutted arithmetic challenges. The slowdown is international and its size is unrelated to countries' consumption or production intensity in information technology — the opposite of what a digital-mismeasurement story predicts. The magnitudes do not reach: existing consumer-surplus estimates for digital technologies are less than one-third of the $2.7 trillion or more of missing output. The implied ICT growth rates are implausible, requiring properly measured ICT output and productivity growth to have been multiples of measured growth. And the gross-domestic-income divergence does not fit: it began before the slowdown and reflects unusually high capital income rather than labour income. Established The first of those is close to decisive against a digital-mismeasurement story specifically.
Frontier “The J-curve and the mismeasurement rebuttal are the same argument.” They are not, and the same economist is a co-author on both. The J-curve is about the timing of intangible recognition; the four challenges are about the level of consumer-surplus omission. Established They are compatible, they are routinely cited as if they were one position, and reading them as one produces the false impression that the mismeasurement case has been simultaneously made and refuted by the same person.
Handwave “It is early.” True in 2017, when the paradox was named and implementation lags were argued to be the largest contributor. Established Nine years later adoption is high, the null is being reported by adopters, and the claim is compatible with every observation available — which is the definition of a handwave rather than a hypothesis. Frontier It becomes testable around 2028–2030, and the tests are already specified: flat aggregate productivity with assisted hours above half retires it, and assisted hours staying in single digits while adoption rises means adoption was shallow rather than early.
Frontier “Self-reported time savings measure productivity.” They measure a belief about time saved. No fetched source validates them against output. Speculative The uncomfortable version of the objection — that experienced practitioners are slower with AI assistance while believing themselves faster — could not be verified in this pass and is stated as a hypothesis with a named consequence: if true, every self-report in this literature is biased upward exactly where use is most sophisticated. Handwave Treating a self-reported saving as an output measurement is the step the argument does by assertion, and it is done in almost every executive survey including the ones cited here.
Speculative “Free digital goods mean GDP is the wrong measure and the gains are hidden in welfare.” Coherent, and the welfare-based measurement work exists: including one large social network alone would have added 0.05 to 0.11 percentage points to annual US welfare-adjusted growth. Frontier That is small relative to the gap it is invoked to close, which is the arithmetic objection rather than a conceptual one. Established The fringe reading is not obviously wrong; it is insufficient by roughly an order of magnitude, and that is a stronger criticism than calling it wrong.
Frontier And the framing itself: “measurably” is carrying weight that nobody has checked. The strongest single fact in this brief is not a null and not an effect size. It is that nine in ten executives at firms actively using AI could not detect a change in their own operations over three years, while the same people forecast one for the next three. Established At the level where productivity is actually produced, the people running the firms cannot measure it. Speculative That is compatible with the gains being absent, being real but small, being real but downstream-blocked, or being real and invisible to the instruments — and this brief declines to pick, because the evidence for picking does not exist yet and the date on which it will is written down.