1 · Concept overview
Established Almost every argument about the future of education runs through a single number. Bloom’s two sigma — the claim that one-to-one tutoring lifts the average student above 98 per cent of a conventionally taught control class — is the target that justified programmed instruction, intelligent tutoring systems, one laptop per child, adaptive courseware, and now the large language model tutor. It is quoted as a finding. The paper that contains it is titled “The 2 Sigma Problem”, its argument is that tutoring is unaffordable at population scale, its evidence is two University of Chicago doctoral dissertations, and its design had three arms, of which the retelling drops the middle one. Bloom was describing a target nobody had hit and asking for help hitting it. Forty-two years later, nobody has hit it.
Established The second thing a reader needs is the deflator, and it appears independently in two separate literatures that the enthusiasts themselves cite. Intelligent tutoring systems measure at 0.73 standard deviations on tests written for the study and 0.13 on standardised tests — a factor of 5.6 from the choice of instrument alone. Mastery learning shows the identical collapse: dramatic effects that, in the words of its own meta-analysts, essentially disappeared when standardised tests were used. Both facts are printed in the reviews everyone quotes. Only the headline numbers travel.
Frontier And the most-cited quantitative claim that generative AI improves learning — a meta-analysis of fifty-one studies reporting a large positive effect — was retracted on 22 April 2026, after independent re-analysis showed the effect did not survive adjustment for publication bias.
This brief takes those three facts as its structure. It covers mastery learning, intelligent tutoring, the personalisation claim, the LLM-tutor trials, assessment and credentialing, the large edtech field experiments, and the exotic question underneath all of them: whether education is information-limited at all, and what the answer implies about which future systems are worth building. Related material sits in Human-AI Integration and Intelligence Measurement.
2 · Current scientific position
Established What Bloom actually wrote. The article is Benjamin S. Bloom, “The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring”, Educational Researcher 13(6), 4–16, June 1984. Bloom’s reported claim is that the average tutored student ended above 98 per cent of the students in the control class, and that about 90 per cent of tutored students reached the summative achievement level reached by only the highest 20 per cent of the control. Established The evidentiary base is two doctoral dissertations completed at the University of Chicago: Joanne Anania (1981) and Arthur Joseph Burke (August 1983). Established The design was three-armed — conventional instruction, mastery learning, and tutoring — so the two sigma is tutoring plus mastery learning against conventional teaching, and the mastery-learning arm alone accounts for roughly the first sigma. Established The title names the finding as a problem, and the problem Bloom poses is that tutoring cannot be afforded at scale, which is why he asks for group methods that match it. Speculative This brief could not obtain the Educational Researcher PDF. The sample sizes, grade levels, subject matter, durations and test types of Anania and Burke are therefore unknown to it, as is Bloom’s full table of alterable variables; it prints none of them. Retrieving that one document would settle more about this subject than any new trial.
Established Mastery learning has its own record and it contains the deflator. Kulik, Kulik and Bangert-Drowns reviewed 108 studies across elementary, secondary and post-secondary levels in 1990 and reported an average effect around 0.59, with mastery programmes most effective for weaker students. Established The same review states the mechanism plainly: by using tests designed for the experiment, mastery instruction may have been able to tailor the class’s learning goals to the measurement tool, and those dramatic effect sizes essentially disappeared when standardised tests were used. Frontier That sentence is doing more work than any result in this brief, and it is forty years old. Established Note also what it implies about Bloom: the middle arm of his own design is the one whose modern meta-analysis is most explicitly instrument-dependent.
Established Intelligent tutoring systems reproduce the collapse independently. Kulik and Fletcher’s meta-analytic review in Review of Educational Research 86(1), 42–78 (2016) found that students receiving intelligent tutoring outperformed conventional classes in 46 of 50 controlled evaluations — 92 per cent, a nearly unanimous direction of effect. Established The median effect size across those 50 studies was 0.66. Established And the average effect on studies using local tests was 0.73, against 0.13 on studies using standardised tests. Established The authors’ own stated conclusion is that alignment of test and instructional objectives is a critical determinant of evaluation results. Frontier Two literatures, developed by different people for different interventions, arrive at the same finding: the instrument is a larger moderator than the intervention. A 5.6-fold swing from the choice of test is not a caveat. It is the main result.
Frontier The field’s own correction to Bloom is VanLehn’s review, and this brief could not read it. Kurt VanLehn, “The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems”, Educational Psychologist 46(4), 197–221 (2011), is the paper the field treats as the deflation of two sigma. Its qualitative headline, as reported in a secondary source, is that there was no statistical difference in effect size between expert one-on-one human tutors and step-based intelligent tutoring systems. Speculative A specific effect size for human tutoring — far below Bloom’s two sigma — circulates widely and is attributed to this review. This brief does not print that figure, because the text could not be obtained through any available route, and the number is the load-bearing element of the entire correction rather than a decoration on it. A brief that repeats an unread number to strengthen its own argument is doing what it accuses the two-sigma literature of doing. Reading pages 197 to 221 would settle it in an afternoon.
Established Real tutoring, under randomisation, measures at roughly three tenths of a standard deviation. Nickow, Oreopoulos and Quan meta-analysed the experimental evidence on PreK–12 tutoring. The 2020 NBER working paper, titled “The Impressive Effects of Tutoring”, reported a pooled estimate of 0.37 SD. Established The peer-reviewed version in the American Educational Research Journal (2024) reports a pooled effect of 0.288 SD — and is titled “The Promise of Tutoring”. Frontier Two things changed between working paper and journal: the estimate fell 22 per cent, and the adjective went. This is peer review working correctly, and the working-paper number is the one that circulates. Established This is intensive, well-run, in-person human tutoring — the intervention Bloom’s two sigma is supposed to describe — measured under randomisation across a systematic review.
Established The magnitude gap, stated carefully, because these are different quantities. Bloom’s target is a claimed effect for tutoring plus mastery learning against conventional instruction, on the underlying studies’ own measures: 2.0 sigma. Intelligent tutoring systems, the technology built to deliver it, measure 0.13 on tests they did not shape. Intensive human tutoring, under randomisation and peer review, measures 0.288 SD. Established So: the target that has driven fifty years of educational-technology investment stands in a ratio of roughly 15 to 1 against what the resulting technology actually measures on independent instruments, and roughly 7 to 1 against what the intervention it describes measures under randomisation. Frontier Both ratios are sourced. Neither is a like-for-like comparison, and this brief does not present them as one: the first compares an aligned-instrument claim against an independent-instrument measurement, which is a category difference as well as a magnitude difference. Both are devastating without exaggeration.
Established The historical record is the field’s least-quoted result. PLATO began at the University of Illinois in 1960 and had NSF funding from 1967. By early 1976 PLATO IV ran 950 terminals with more than 3,500 contact hours of courseware; the library eventually exceeded 12,000 contact hours, the largest ever built. External evaluation found it essentially equal to an average human teacher in terms of student advancement, with the observation that everyone using it nevertheless enjoyed it. Established The economics killed it: terminals around $12,000, Control Data charging $50 an hour for data-centre access, and courseware development averaging $300,000 per delivery hour. Established Teaching machines go back further — Pressey in the mid-1920s, Skinner in the 1950s and 1960s — and the fetched account of that literature contains the sentence that ought to be on the wall of every edtech company: there is extensive experience that both methods worked well, and so did programmed learning in other forms, such as books. The machine was not the active ingredient.
Established One Laptop Per Child is the largest natural experiment ever run on technology-led education reform, and it is null. A randomised evaluation in Peru covering 531 primary schools over ten years found no significant effects on academic performance, on primary or secondary completion, or on university enrolment. Established Uruguay’s Plan Ceibal, evaluated in 2013 by the Economics Institute of the University of the Republic, found no impact on reading or mathematics test scores — and recorded that only 4.1 per cent of laptops were used all or most days in 2012. Frontier That usage figure is the mechanism showing itself, and it is the single most under-reported number in educational technology. Established More than three million laptops shipped by 2015; the $100 price target was never met, the price standing above $209 in April 2011. Frontier The record is contested at the edges — the same source notes more recent studies regarding the project as a success — but the Peruvian trial is the largest randomised evidence available and it found nothing.
Established Personalisation’s one hard test came back null, and belief survived it. Pashler, McDaniel, Rohrer and Bjork concluded in Psychological Science in the Public Interest in 2008 that there is no adequate evidence base to justify incorporating learning-styles assessments into general educational practice; that studies using the crossover design necessary to test the claim were virtually absent from the literature; and that among those with proper methodology, all but one produced negative findings. Coffield and colleagues had catalogued 71 distinct models. Established A 2017 United Kingdom survey found 90 per cent of academics agreeing that learning-styles theory has basic conceptual flaws, 58 per cent nonetheless agreeing that students learn better when information matches their preferred style, and 33 per cent having used learning styles in the past year. Frontier That gap is the most important single fact for assessing AI-personalisation claims: the demand for personalisation is demonstrably not evidence-driven, so evidence against a particular personalisation will not reduce it.
Frontier The LLM-tutor evidence, in full, is four studies pointing in three directions. Kestin and colleagues in Scientific Reports (2025) ran a crossover trial in Harvard’s Physical Sciences 2 course: 194 eligible students, two lessons in consecutive weeks, AI-tutor median post-score 4.5 against 3.5 for in-class active learning, linear-regression effect 0.63, quantile regression between 0.73 and 1.3 SD, and a median 49 minutes on task against roughly 60 in class. Engagement and motivation were both higher in the AI condition. Frontier Bastani and colleagues in PNAS (2025) ran a field experiment with around a thousand high-school mathematics students and reported that when unguarded access was subsequently taken away, students performed worse than those who never had access — a 17 per cent reduction in grades for the unguarded arm. Frontier Wang, Ribeiro, Robinson, Loeb and Demszky’s Tutor CoPilot trial — 900 tutors and 1,800 K-12 students from underserved communities, the first randomised trial of a human-AI tutoring system — found students of participating tutors 4 percentage points more likely to master topics, rising to 9 points for students of the lowest-rated tutors, at $20 per tutor per year. Frontier And a 2026 randomised trial of LLM-driven Socratic questioning in endodontic training found no statistically significant difference from faculty instruction while substantially reducing faculty time. Established Read together, these are not a consensus. They are an open question with four good measurements in it.
3 · Frontier questions
Established The frontier event in this subject is a retraction. Wang and Fan’s meta-analysis of ChatGPT’s effect on learning — 51 studies, learning performance g = 0.867, learning perception 0.456, higher-order thinking 0.457 — was published in Humanities and Social Sciences Communications on 6 May 2025 and retracted on 22 April 2026. For eleven months it was the most-cited quantitative claim that generative AI improves learning. Any brief, pitch deck or policy paper still quoting g = 0.867 is quoting a retracted paper.
Frontier The mechanism of the correction is publication-bias adjustment, and it is the methodological frontier for all of edtech. Bartos, Martinkova and Wagenmakers showed that these effects greatly diminish once publication bias is accounted for and that the evidence in favour of the benefits disappears; they call for high-quality pre-registered experiments. Frontier A second, independent re-analysis by Jafri put higher-order thinking at g = 0.309 with a 95 per cent interval of −0.021 to 0.640 — statistically indistinguishable from zero. Frontier Two teams, arriving separately, at the same place. This is not a dispute about one paper; it is a diagnosis of a literature.
Frontier The withdrawal effect is genuinely new and has no analogue in the historical record. No PLATO, OLPC or intelligent-tutoring study measured what happens to a learner when the system is removed, because none of those systems could do the work for the student. Frontier Bastani and colleagues did measure it, and the direction is negative. Speculative A published critique of that design by Tan and Rajaratnam identifies potential confounding factors, and the paper’s own structure includes a guarded tutor arm that did not show the harm — which is why the peer-reviewed title is conditional. Frontier This is the education-specific instance of the deskilling and automation-dependence findings assembled in Human-AI Integration, and it inverts the policy question from “does it help while in use” to “what state is the learner in when it is taken away”.
Frontier The augment-the-tutor design is outperforming the replace-the-tutor design on study quality, not just on results. Tutor CoPilot is the largest randomised trial in the set, targets the human, analysed over 550,000 messages, and found participating tutors more likely to use high-quality strategies such as asking guiding questions. Kestin targets the student and is 194 students in one course. Bastani targets the student and found harm. Speculative That is a pattern rather than a proof, and it matches the cross-domain generalisation that the robust positives in human-AI work lift the bottom of a skill distribution: Tutor CoPilot’s largest gains were for the lowest-rated tutors, exactly as mastery learning’s largest gains were for the weakest students.
Frontier Non-inferiority at radically lower cost is emerging as the defensible claim, and it is a different claim from personalisation. PLATO was found essentially equal to an average human teacher. The endodontic trial found no difference at substantially reduced faculty time. Tutor CoPilot costs $20 per tutor per year. Speculative If the honest finding is “same learning, far less expert time”, that is an enormous result about access and cost, and it should be pursued and reported as such. Frontier The field keeps claiming superiority instead, and the superiority claims keep failing bias adjustment.
Frontier The measurement problem Kulik and Fletcher named in 2016 is about to get worse. An LLM tutor optimised against a course’s own assessments is the strongest possible version of the alignment confound, because the system can be tuned toward the instrument continuously and at no cost. Established Kestin’s outcome was a purpose-designed post-test. Frontier Every LLM-tutor study should be read first for what test it used and only second for its effect size. Speculative The corresponding hazard on the assessment side is contamination: a public benchmark or standard examination that the model has seen is not a measurement of the student.
Frontier The field is responding correctly, and the 2027–2029 evidence will be much better than the 2023–2025 evidence. Two AI-tutor randomised trials were entered in the AEA RCT Registry in 2025 alone — conversational bots teaching microeconomics as scalable instructional supplements, and an AI tutor for Colombian undergraduates in post-secondary programming, in asynchronous and blended formats. Speculative Pre-registration is the specific remedy for the failure mode that produced the retraction, and its adoption in this subject is recent enough to date precisely.
Frontier The largest deployed system has the thinnest published evidence. Khanmigo launched on 14 March 2023 on GPT-4 at a $4 monthly paid tier. The published evidence for it, as of this writing, consists of a February 2024 Wall Street Journal test finding basic calculation errors and a 2026 PNAS evaluation of Khan Academy’s effect on mathematics learning whose effect sizes this brief could not obtain. Frontier Khan Academy’s own headline — twenty hours of practice associated with a 115-point SAT increase — is stated as an association, and hours-studied is the variable most confounded by motivation and prior attainment.
Speculative The cognitive-debt question is open and badly measured. Kosmyna and colleagues’ essay-writing study argues for an accumulation of cognitive debt under LLM assistance. It is a small, single-laboratory, preprint result with EEG endpoints of contested interpretability, and it has been reported far beyond what its design supports. Frontier It nonetheless asks the right question, which is the same one Bastani’s withdrawal arm asks with a better design and a harder outcome.
4 · Technological bottlenecks
Established The binding constraints in this subject are measurement practices, not capabilities. That is unusual, and it is the most useful thing this brief can say. Nothing below requires a new model, a new device, or a large budget.
Frontier Bottleneck one: nobody has decomposed the tutoring effect. A tutor supplies information, pacing and mastery gating, accountability and attendance, and a relationship. No study separates them, so no one knows which components a machine can supply. Established Every AI-tutor trial ever run is a bet on an unmeasured decomposition. Speculative If the effect is mostly accountability — a specific adult watching a specific child do specific work at a specific time — then AI tutoring captures the small term and the ceiling is low regardless of model capability.
Frontier Bottleneck two: adaptivity has never been isolated. Kulik and Fletcher’s 0.66 compares intelligent tutoring against conventional classes, which confounds adaptivity with content quality, feedback, and time on task. Established The entire premise of “AI will personalise education” rests on a comparison nobody has run: adaptive sequencing against a fixed expert-designed sequence, content and time held constant. Frontier The learning-styles precedent is the cautionary case — 71 models, an enormous literature, and the crossover design that would have tested it virtually absent.
Established Bottleneck three, and historically the killer: instrument independence. 0.73 to 0.13 for intelligent tutoring; effects that essentially disappeared for mastery learning. Frontier Any result that has not been shown to survive on a measure the system did not shape is not yet a result about learning.
Frontier Bottleneck four: survival of withdrawal. The only measurement anyone has of what happens after access ends is negative.
Established Bottleneck five: the thing has to get used. 4.1 per cent of Plan Ceibal laptops in daily use is the modal failure mode of educational technology, and it is almost never a pre-registered endpoint. Frontier A dosage-response curve is worthless without a usage distribution beside it.
Speculative The one bottleneck that has plausibly moved is cost. PLATO’s $300,000 per delivery hour of courseware is what content authoring used to cost. Generative models change that number by orders of magnitude, and that is the strongest available reason to think this cycle is not the seventh repetition of the previous six.
5 · Research dependencies
Established This brief waits on one result another brief produces. Whether human-AI combination beats the better of its components is the general question of which AI tutoring is a special case, and the meta-analytic answer assembled in Human-AI Integration is that on average it does not — with the important qualification that the robust positives augment the weaker party toward an existing standard. Frontier Tutor CoPilot is exactly that shape and it works; the replace-the-tutor designs are the other shape and they are contested.
Frontier The deeper dependency is on the science of measuring learning itself. A field whose effect sizes move by a factor of 5.6 on instrument choice does not have a settled outcome variable. Progress here depends on assessment validity work that sits closer to Intelligence Measurement than to educational technology, and specifically on instruments with known transfer properties.
Established It depends on cognitive load theory for its content design. The worked-example effect and the expertise-reversal effect are the best-established design constraints in instruction, and they say that the optimal presentation for a novice is actively harmful for an expert — which is a real, measured, mechanism-bearing form of personalisation that predates and outperforms the learning-styles version.
Speculative And it depends on a decomposition nobody owns. The tutoring-components question falls between education research, experimental economics and psychology, and no funder treats it as a priority, which is why a cheap and decisive experiment has gone unrun for four decades. Frontier The pattern generalises: the results this subject most needs are the ones that would deflate a claim rather than support a product, and the deflating result has no natural sponsor.
Frontier One dependency is negative and worth stating. Nothing here waits on a larger model. Every binding link in the workback plan below is satisfiable with the systems that already exist, and none of the historical failures — PLATO, programmed instruction, one laptop per child — failed for want of capability. Speculative If that is right, capability progress will change the cost of this subject substantially and its outcomes very little, which is an uncomfortable prediction and a testable one.
6 · Required experiments
Established The workback plan, ordered, with the binding links named. Set the target honestly first: on standardised measures, the best rigorous estimate for intensive human tutoring is 0.288 SD. A programme aimed at two sigma is aimed at a number never reproduced on an independent instrument. Aim at 0.3 SD population-wide on standardised measures — still far beyond anything any system has achieved — and the programme becomes tractable.
Frontier Experiment 1 — replicate Bloom faithfully. A pre-registered, powered, three-arm trial: conventional instruction, mastery learning, mastery learning plus tutoring. Report a standardised outcome and a local outcome side by side; the gap between them is itself a headline result. Established It does not bind — it is cheap — but it determines whether the target is 2.0 or 0.3, which determines everything downstream. It is astonishing that it has not been done.
Frontier Experiment 2 — decompose the tutoring effect. Binding. An additive-factorial design separating informational content, pacing and mastery gating, accountability and attendance, and relational or affective support, with per-component effect sizes. Until this exists nobody knows which component a machine can supply.
Frontier Experiment 3 — isolate adaptivity. Binding. Adaptive sequencing against a fixed expert-designed sequence, content and time held constant, standardised outcome. This is the entire premise of AI personalisation and it has never been cleanly tested.
Established Experiment 4 — dual-instrument reporting. Binding, and the highest-value action available. Every study reports effect size on both a local and an independently constructed standardised measure. This is a reporting norm rather than an experiment; it costs one extra assessment; and it would have prevented the two-sigma myth, most of the intelligent-tutoring literature’s overclaiming, and the retracted meta-analysis.
Frontier Experiment 5 — survival of withdrawal. Binding. Bastani’s design extended, with the guardrail arm, and post-access assessment showing treated students at or above never-treated controls.
Frontier Experiment 6 — population scale outside elite institutions. Binding. Kestin-magnitude effects in a non-selective population, over a term, multi-site cluster-randomised, standardised outcomes. OLPC in Peru is the cautionary case at 531 schools.
Established Experiment 7 — usage as a pre-registered endpoint. Binding, and the one that killed OLPC. Telemetry reported against the trial’s own dosage-response curve.
Speculative Experiment 8 — cost per standard deviation at scale, including content development. Probably no longer binding; this is the link generative models plausibly break.
7 · Engineering requirements
Established Mastery gating is the engineering requirement the retelling of Bloom deleted. Bloom’s design gated progression on demonstrated mastery with corrective loops and flexible time-to-mastery. Almost nothing built since has implemented that half properly, because it conflicts with fixed timetables, fixed cohorts and fixed term lengths. Frontier A system that personalises explanation but not pacing has implemented the arm that measures smaller.
Established Content authoring cost is the variable that has genuinely changed. PLATO’s courseware averaged $300,000 per delivery hour, against a library of 12,000 contact hours — the reason it failed commercially despite good evaluations. Speculative Generation collapses that cost, but it moves the bottleneck to verification: the Wall Street Journal finding that a deployed tutor made basic calculation errors is what unverified generated instruction looks like at scale.
Frontier Assessment generation is where the system can silently mark its own homework. If the same model authors the instruction and the assessment, the local-test effect size becomes uninterpretable by construction, and the 0.73-versus-0.13 gap becomes unmeasurable rather than merely unreported. Speculative The engineering answer is an assessment supply chain independent of the instruction supply chain, with items the instructional system has never seen. Nobody ships this.
Frontier Guardrails are now a specified component rather than a safety afterthought. The peer-reviewed claim in the strongest negative study is conditional on their absence, and the guarded tutor arm did not show the harm. Speculative What a guardrail is — refusing to produce answers, requiring an attempt first, Socratic constraint — is under-specified, and the difference between arms in that trial is currently the field’s only calibration for it.
Established And usage telemetry is a first-class requirement, not instrumentation. The 4.1 per cent figure was recoverable only because someone measured it, and it is the difference between a system that works and a system that is installed. Frontier A deployment that cannot report the distribution of minutes per learner per week cannot support any claim about its own effect, because the dosage-response curve is unanchored. Speculative The corresponding design requirement is that the system degrade gracefully at low usage rather than assuming the modal learner is the enthusiastic one, which is the assumption every courseware library since PLATO has quietly made.
8 · Adjacent technologies
Established The best-evidenced intervention in this entire subject is free and involves no technology. The testing effect — that retrieval practice produces more durable learning than restudy — is robust, replicated, and demonstrated in classrooms rather than only in laboratories. Frontier This brief does not print a pooled effect size for it, because the meta-analytic values were not obtainable here; the qualitative result is not in dispute. Speculative A system whose entire contribution was to schedule retrieval practice and spacing correctly would be a more defensible product than most of what is sold as adaptive learning.
Established Cognitive load theory supplies the one form of personalisation with a mechanism. The worked-example effect and its expertise reversal say that the presentation optimal for a novice degrades performance for an expert. That is an interaction between instruction and learner state that has been measured repeatedly, unlike the learning-styles interaction, which has not.
Frontier Automated assessment is the adjacent technology with the largest institutional consequences. When the assessment can be produced and completed by the same class of system, the assessment stops measuring the student and starts measuring access. Speculative The plausible responses — oral examination, in-person invigilation, process evidence, portfolio defence — are all more expensive per student than the assessments they replace, which points the cost curve in the opposite direction from the tutoring one.
Frontier Two further adjacencies. The human-AI complementarity literature in Human-AI Integration supplies the ceiling on any “teacher plus AI” claim. And the individual-capability question — whether a learner can be made to learn faster at all — belongs to Cognitive Enhancement, whose answer is that the honest measured effects there sit between 0.1 and 0.3 SD, in the same band as everything in this brief. Speculative Two fields with different mechanisms, different literatures and different failure modes converging on the same effect-size band is either a deep fact about how much a healthy human’s learning can be moved by any intervention at all, or a fact about how psychology and education measure things. Nobody has argued it either way in print.
9 · Institutional requirements
Established The cheapest institutional reform available in this subject is a reporting rule. Require every education-effect study to report both a locally constructed and an independently constructed standardised outcome. It costs one assessment. It is within the gift of any journal editor, any funder, and any ministry procurement office. Frontier On the historical record it would have prevented the two-sigma myth from hardening, would have deflated much of the intelligent-tutoring literature at the time rather than in 2016, and would have flagged the retracted meta-analysis before publication.
Frontier Pre-registration is the second, and it is already arriving. Two AI-tutor trials entered the AEA RCT Registry in 2025. Speculative A funder that made registration a condition of grant for education-technology evaluation would change the 2029 evidence base more than any product decision made this year.
Speculative The sorting hypothesis is the strongest institutional argument that the constraint is not technological. If credentials function substantially as positional signals, then an intervention that raises everyone’s learning leaves relative position unchanged and destroys the signal’s value; institutions would then resist adoption for reasons uncorrelated with efficacy. Speculative The evidence is circumstantial: OLPC shipped three million laptops and produced no systemic change; PLATO had the best courseware library ever built and failed commercially. Handwave The clean falsifier — a learning technology adopted at scale specifically because it compressed outcome variance — has no instance anyone can name, which is either strong support or an artefact of nobody looking.
Frontier Credentialing and learning are separable and are being separated by force. If assessment can be automated on both sides, the credential’s informational content falls, and institutions must either raise assessment cost per student or accept a weaker signal. Speculative That is a budget decision disguised as a pedagogical one, and it will be made in procurement rather than in research.
Established Teacher labour is the actual delivery constraint. The one modern result with a defensible cost story augmented tutors rather than replacing them, at $20 per tutor per year, and helped the weakest tutors most. That is an institutional design, not a model capability.
10 · Ethical & societal considerations
Frontier The withdrawal finding creates an obligation that does not exist elsewhere in edtech. If unguarded access leaves a student worse off than never having had access, then deployment without a plan for what happens when access ends — a school year, a subscription, a funding cycle — is a foreseeable harm. Speculative No procurement framework currently asks the question.
Established The subjects are children, and the trials are run in schools. A field experiment on roughly a thousand high-school mathematics students that measured a grade reduction is an ordinary and necessary piece of research; it is also a reminder that the population being experimented on cannot decline, and that the consent structure is institutional rather than individual.
Frontier The equity case is real and it runs through variance reduction. Mastery learning’s largest gains were for weaker students; Tutor CoPilot’s largest gains were for students of the lowest-rated tutors. Speculative If that pattern is the reliable one, the honest ethical claim for AI in education is compression of the outcome distribution at low cost, not elevation of the mean — and it is a better claim than the one being made.
Speculative The relational objection deserves a straight answer. If part of what a tutor supplies is being cared about by a person, a simulation that is believed is a deception and one that is disbelieved is inert. Frontier The only direct evidence points the other way: in the Harvard trial, engagement and motivation were both higher in the AI condition than in the human-taught one. That is short-term, in adults, with grades attached, and it should not be over-read — but it is what the measurement says.
Frontier And the measurement problem is an ethical problem. Selling a 0.73 local-test effect to a ministry that will be judged on standardised results is not a technical disagreement about instruments.
11 · Civilizational implications
Speculative The central exotic hypothesis of this subject is that education has not been information-limited since roughly the printing press. Access to high-quality explanation has been effectively free for decades — libraries, textbooks, PLATO’s 12,000 contact hours, Khan Academy, video. Population learning outcomes did not move. If that is right, the binding constraints are attention, motivation, prior knowledge, time on task, assessment structure, and the social contract that induces a young person to do difficult things they would rather not do — and every technology whose value proposition is cheaper access to explanation is aimed at a non-constraint.
Frontier The evidence for it is the strongest negative evidence in the brief. 531 schools and ten years with no effect; 4.1 per cent daily usage; the largest courseware library ever built evaluating as equal to an average teacher; programmed learning working equally well as books. Frontier The evidence against it is that tutoring does work at 0.288 SD, and part of what a tutor does is informational. Speculative Nobody has decomposed it, which is why the decomposition experiment is the binding link.
Speculative If the hypothesis holds, the civilizational implication inverts the investment case. The lever is not explanation supply but the structures that produce compliance, attention and time on task — which are institutional, social and political objects, not products. Handwave A civilisation that solved education in that reading would look less like better software and more like a different arrangement of childhood, adult time and social obligation, and no one has a workback plan for that.
Speculative And if it fails — if the informational component turns out to dominate after all — then the cost collapse in content generation is the most consequential development in the history of instruction, and this brief will have been wrong in the most useful possible direction.
12 · Timelines
These horizons track results, not products. The dates below are dates by which a specific measurement either exists or does not.
- 10 yr: Frontier The pre-registered AI-tutor trials now in the AEA registry report, and dual-instrument reporting either becomes normal in education journals or does not; that single editorial decision determines how interpretable the 2030s literature is. Frontier The withdrawal effect is either replicated with a guardrail arm or fails to replicate, and the answer arrives from a field experiment rather than a laboratory. Speculative At least one Kestin-style result is attempted in a non-selective population over a full term; on the base rate it comes in substantially smaller.
- 25 yr: Speculative The tutoring decomposition has been run, and the field knows for the first time what fraction of the effect is informational. Speculative Adaptivity has been isolated against a fixed expert sequence; the learning-styles precedent says the honest prior is a small effect, not a large one. Frontier Assessment has been rebuilt around what cannot be automated, at higher cost per student, and the credential has partially separated from the course.
- 50 yr: Speculative Either an education system somewhere is delivering 0.3 SD population-wide on instruments it did not shape — which would be the first time in the measured history of the subject — or the constraint has been demonstrated to be institutional and the research programme has moved to attention, time and social structure. Speculative Non-inferiority at a tenth of the cost is the more likely achieved outcome and would transform access without moving the mean much.
- 100 / 250+ yr: Handwave Tutoring-scale gains at population scale require six things to hold at once: a known decomposition, a machine-suppliable dominant component, gains that survive independent instruments, gains that survive withdrawal, gains that survive non-elite populations, and a system people actually use. Handwave Each has failed at least once in the measured record and four have never been tested. Handwave Cheerfully far from evidence, and stated here because the alternative — asserting that scale plus capability will do it — is what the previous six cycles of this technology asserted.
13 · Technology tree & dependencies
- Depends on This brief depends on Human-AI Integration for the general result of which AI tutoring is a special case: whether a human-machine combination outperforms the better of its components, and under what conditions. The meta-analytic answer there is on average negative, with the robust positives concentrated where the system augments the weaker party toward an existing standard — which is precisely the shape of the one large randomised tutoring trial that worked, and precisely not the shape of the replace-the-teacher designs. That is a typed dependency rather than a thematic one: if the complementarity result moves, the prior on every AI-tutoring design in this brief moves with it, and the direction of the move is not currently predictable from anything inside education research. The dependency also runs the other way in evidential weight. Education is the largest single deployment anyone proposes for human-AI pairing, so a tutoring trial at population scale would be among the most informative measurements the complementarity literature could receive, and none of the trials assembled here was designed to serve that purpose.
- Enables A settled decomposition of the tutoring effect would enable everything downstream in this subject, because it would say for the first time which component of a tutor a machine can supply and which it cannot. Nothing on this map currently waits on that decomposition, and that absence is itself a finding rather than an oversight: the binding link in education is a cheap factorial experiment that no adjacent research programme has claimed, because it sits between education research, experimental economics and cognitive psychology and is owned by none of them. The same is true of the instrument reform. A rule requiring both a local and a standardised outcome enables honest reading of every future result in this subject and costs one assessment per trial, and it has no natural owner either, which is why forty years of literature has been readable only by people who already knew to check.
- Adjacent Adjacent work sits in Intelligence Measurement, which owns the instrument problem that makes education effect sizes swing by a factor of 5.6 and which will have to supply any assessment with known transfer properties; in Cognitive Enhancement, whose honest measured effects on healthy adults fall in the same 0.1 to 0.3 band as everything in this brief, which is a coincidence worth noticing rather than explaining away; and in Distributed Cognition, which owns the question underneath the withdrawal result — what is actually lost when a capability is offloaded rather than acquired, and whether the loss is a deficit or a reallocation.
14 · Common misconceptions & speculative claims
Established “Bloom proved one-to-one tutoring produces a two-standard-deviation gain.” Four things are wrong with that sentence. The paper is titled “The 2 Sigma Problem” and its argument is that tutoring is unaffordable at scale, so it is a statement of an unsolved problem cited as a delivered result. The evidence is two University of Chicago doctoral dissertations — Anania (1981) and Burke (1983) — not a programme of trials. The design was three-armed and the intervention was tutoring plus mastery learning against conventional instruction, with the middle arm dropped from essentially every retelling. And the best modern meta-analysis of randomised tutoring evidence reports 0.288 SD. Frontier This is the clearest case on this site of a founding paper being cited for the opposite of what it says.
Established “Two sigma is robust and replicated.” The replication record runs the other way. VanLehn’s 2011 review — the field’s own correction — found no statistical difference between expert human tutors and step-based intelligent tutoring systems, which means human tutoring is not in a class of its own. Kulik and Fletcher’s intelligent-tutoring median is 0.66, falling to 0.13 on standardised tests. Mastery learning’s effects essentially disappeared on standardised tests. Speculative As far as this brief can determine, the two-sigma figure has never been reproduced on an independently constructed assessment. It also could not read VanLehn, and says so rather than borrowing his numbers.
Established “A meta-analysis shows ChatGPT substantially improves learning, g = 0.87.” That meta-analysis was retracted on 22 April 2026. Before the retraction, an independent re-analysis had already shown the effects greatly diminish once publication bias is accounted for and that the evidence in favour of the benefits disappears; a second re-analysis put higher-order thinking at g = 0.309 with a confidence interval spanning zero. Frontier It was published in a Nature-portfolio journal and stood for eleven months. Anyone still citing 0.867 is citing a retracted paper.
Established “Effect sizes in education are comparable across studies.” Intelligent tutoring measures 0.73 on local tests and 0.13 on standardised tests — a 5.6-fold difference produced by the choice of instrument alone, in the same set of studies. Mastery learning shows the identical collapse in a separate literature developed by different people. Both are stated in the reviews the enthusiasts cite; only the headline numbers travel. Frontier Before comparing any two education effect sizes, check what test each used. A number reported without its instrument is not a measurement.
Established “AI will personalise education, and personalisation is what education has always lacked.” The one personalisation hypothesis that has been tested hard — matching instruction to learning style — has no adequate evidence base, with proper-design studies producing all but one negative finding across a literature containing 71 distinct models. Adaptivity as such has essentially never been isolated in an intelligent-tutoring trial. Frontier And the demand for personalisation is demonstrably not evidence-driven: 90 per cent of surveyed academics acknowledged the conceptual flaws while 58 per cent still endorsed the meshing claim and 33 per cent had used learning styles that year. Speculative Pace-and-mastery personalisation may genuinely be different from style-matching. It has not been shown to be.
Frontier “AI tutors have been shown to beat classroom instruction.” One study shows this: 194 Harvard undergraduates, one lesson, crossover design, purpose-built post-test, content authored by physics-education researchers, effect 0.63. A field experiment with about a thousand high-school students found unguarded access produced a 17 per cent grade reduction once withdrawn. A 2026 professional-training trial found no significant difference from faculty instruction. Established Quoting only the first is not describing the literature. Frontier And the pessimists misuse the second exactly as the optimists misuse the first: the peer-reviewed title is conditional on the absence of guardrails, the preprint title was not, and the widely circulated version is the preprint’s.
Established “This time the technology is different.” Pressey’s teaching machine is from the mid-1920s and Skinner’s from the 1950s, and the record states that programmed learning worked well in other forms, such as books — the machine was never the active ingredient. PLATO assembled 12,000 contact hours of courseware and evaluated as essentially equal to an average human teacher. OLPC shipped three million laptops, and Peru’s 531-school, ten-year randomised evaluation found no effect on performance, completion or university enrolment. Speculative The claim may still be true — the content-authoring economics have genuinely changed by orders of magnitude — but it has to be argued against this record rather than asserted over it.
Established “Kids will teach themselves if you give them access.” The Hole in the Wall kiosks, begun in Delhi in 1999 and replicated in Cambodia in 2004, are the standard citation. Trucano found no evidence of increases in the key skills claimed; Clark found the machines dominated by older boys, excluding girls and younger children, and used mostly for entertainment; Arora documented kiosks falling into disrepair and abandonment without the resources typical of a school; Cuban characterised the design as dumping hardware in schools and hoping for magic. Frontier The claim survives because it is charming, not because it was measured.
Established “Khan Academy’s data proves the model works — twenty hours gives 115 SAT points.” That is stated as an association, and hours-studied is precisely the variable most confounded by motivation and prior attainment. The same record notes a Wall Street Journal test finding the deployed AI tutor made basic calculation errors. Frontier A 2026 PNAS evaluation exists and may change this assessment; its effect sizes were not obtainable here and none are printed.
Frontier “Peer review filters this.” Both halves are on the record. The tutoring meta-analysis went from “The Impressive Effects” at 0.37 SD to “The Promise” at 0.288 SD between working paper and journal — peer review working exactly as intended. And a meta-analysis reporting g = 0.867 was published in a Nature-portfolio journal and stood eleven months before retraction — peer review failing. Speculative The asymmetry that matters is which version circulates: the working-paper number and the retracted number both had months of unopposed currency, and neither correction travelled as far as the claim it corrected.
Speculative “Bloom was simply wrong.” This brief does not assert that, and the honest terminal position is a declared tie. It is entirely possible that tutoring plus properly implemented mastery learning does produce very large effects under the right conditions, and that everything since has failed to reproduce the conditions rather than demonstrating the effect was illusory — almost nothing built in forty years implemented the mastery half faithfully. Handwave The falsifier is cheap, obvious and unrun: a pre-registered three-arm replication with standardised outcomes. The field has spent four decades citing an experiment it never repeated, and the reason to run it is not that Bloom was wrong but that nobody knows.