1 · Concept overview

Established Two different claims travel under the name artificial general intelligence, and almost every public argument about it conflates them. The first is a claim about achievability: that a system with broad, transferable competence across domains it was not specifically built for is reachable by some path from here. The second is a claim about recognition: that if such a system arrived, we would be able to tell. Frontier This brief treats the second as the harder of the two, and treats the gap between them as the field's central unresolved problem rather than as a technicality. Established Every headline capability number in this subject is a number about a protocol — which evaluation split, how many attempts, how much inference compute, and whether the test existed publicly before the system was trained — and a large fraction of the field's most-quoted figures collapse or reverse when the protocol is attached.

Established The field cannot agree on a definition because two operationalizations are in circulation that are not reconcilable. Chollet defines intelligence as skill-acquisition efficiency relative to priors and experience, so that a system trained on ten million examples of a task family has demonstrated nothing by solving that family. Established Morris et al. define it as depth of performance crossed with breadth of domains, so that the same system has demonstrated general capability. Frontier These give opposite verdicts on the same evidence, and no result in the literature adjudicates between them. A brief on this subject either picks a side or makes the fork its subject; this one makes the fork its subject.

Frontier The strongest honest statement of the current position is a shape rather than a threshold. The best-resourced attempt to date to build a psychometric definition reports a profile with superhuman performance in some faculties and severe deficits in others, which is not the sort of object a yes/no question has an answer about. Speculative The exotic reading, taken seriously here, is that “general intelligence” is a folk concept imported from a human population with a shared developmental architecture, and that it may not carve machine systems at any joint at all. Established The measurement problem underlying all of it belongs to Intelligence Measurement, on which this brief depends.

2 · Current scientific position

Established The scaling laws claim considerably less than they are cited for, and the original prescription was overturned within two years. Kaplan et al., arXiv:2001.08361v1 (submitted 2020-01-23), report that language-model loss “scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude.” Established That is a claim about loss, not about capability, and the paper's load-bearing prescription — that “optimally compute-efficient training involves training very large models on a relatively modest amount of data” — was contradicted by Hoffmann et al., arXiv:2203.15556v1 (2022-03-29), who trained over 400 models from 70M to 16B+ parameters and concluded that “for compute-optimal training, the model size and the number of training tokens should be scaled equally.” Established Chinchilla at 70B parameters and four times Gopher's data outperformed Gopher at 280B and reached 67.5% on MMLU. Frontier Nothing in either paper licenses an inference from a loss curve to a capability, and the step from one to the other is where most public scaling arguments quietly do their work.

Established The emergence literature has been through a full cycle of claim and deflation, and the deflation is now the better-supported side. Wei et al., arXiv:2206.07682 (TMLR 2022), defined an emergent ability operationally as one “not present in smaller models but is present in larger models,” and therefore not predictable by extrapolation. Established Schaeffer, Miranda and Koyejo, arXiv:2304.15004, answered that “nonlinear or discontinuous metrics produce apparent emergent abilities, whereas linear or continuous metrics produce smooth, continuous predictable changes,” supporting it with predictions across the GPT-3 family, a BIG-Bench meta-analysis, and an induced-emergence demonstration on vision models. Established BIG-Bench itself, 204 tasks and 450 authors across 132 institutions (arXiv:2206.04615), found that tasks showing gradual improvement typically involve memorization while tasks showing breakthrough behaviour at critical scale “frequently involve multiple steps or unreliable metrics,” and that social bias typically increases with scale in ambiguous contexts. Frontier The correct residual claim is narrow: apparent discontinuities are largely metric artifacts, which is not the same as showing that no capability ever appears discontinuously.

Established ARC-AGI is the field's flagship generalisation test, and its own numbers are the sharpest available demonstration of why a score without a protocol is not a fact about the world. Chollet's founding argument, arXiv:1911.01547, is that “solely measuring skill at any given task falls short of measuring intelligence, because skill is heavily modulated by prior knowledge and experience: unlimited priors or unlimited training data allow experimenters to buy arbitrary levels of skills for a system, in a way that masks the system's own generalization power.” Established ARC Prize 2024 (arXiv:2412.04604) records that state of the art on the ARC-AGI private evaluation set rose from 33% to 55.5% during 2024, against a prize target of 85%. Established The ARC-AGI-2 paper (arXiv:2505.11831) reports, in Table 1 on the semi-private set: the 2024 competition-winning ARChitects system at 56.0% on ARC-AGI-1 and 2.5% on ARC-AGI-2; o3 (Medium) at 53.0% and 3.0%; o4-mini (Medium) at 41.8% and 2.4%; o3-mini (High) at 34.5% and 3.0%; Claude 3.7 (8K) at 21.2% and 0.9%. Human baselines in the same paper: 407 unique participants across 515 sessions, 8,277 of 13,405 test-pair attempts solved (62%), 75% of tasks completed when aggregated by task, and for ARC-AGI-1 the observation that people at the higher end of the distribution “could solve over 97% of ARC-AGI-1 tasks without much effort.” Frontier The gap between roughly 3% and roughly 75% on ARC-AGI-2 is the single largest human-machine discrepancy on any current benchmark, and it appeared immediately after the previous version had been driven to within thirty points of its prize target.

Established The December 2024 o3 that was demonstrated on ARC-AGI is not the o3 that shipped, and the ARC-AGI-2 paper documents both inside one document. That paper's body reports the o3 preview achieving “76% (low-compute; estimated cost: $200 per task)” and “88% (high-compute; estimated cost: $20,000 per task)” on the ARC-AGI-1 semi-private set. Established Its Table 1 reports the released o3 (Medium) at 53.0% on that same set. Established That is a 23-to-35 point gap on the field's most-quoted progress number, on the same evaluation split, in the same paper. Frontier It is not evidence of dishonesty by anyone; it is evidence that “system X achieves Y%” is an incomplete sentence in this field, and that the missing arguments — which build, at what inference cost, on which split — can be worth more than a factor of one and a half. Established A second detail belongs with it: the widely circulated pair “75.7% / 87.5%” does not appear in the ARC-AGI-2 paper, which reports 76% and 88%; the original ARC Prize announcement could not be retrieved for this brief, so the more precise pair is quoted here as unsourced rather than repeated as fact.

Established Long-horizon agentic evaluation is the most methodologically serious measurement programme in the subject, and it is scoped to software by its own authors. METR's time-horizon paper proposes the “50%-task-completion time horizon” — the time humans typically take on tasks a model completes with 50% success. Established From arXiv:2503.14499v2, the version this brief was written from: models were evaluated on “a combination of RE-Bench, HCAST, and 66 novel shorter tasks”; “current frontier AI models such as Claude 3.7 Sonnet have a 50% time horizon of around 50 minutes”; the horizon has been “doubling approximately every seven months since 2019, though the trend may have accelerated in 2024”; and the increase is “primarily driven by greater reliability and ability to adapt to mistakes.” Established The extrapolation is stated conditionally in the paper itself: “If these results generalize to real-world software tasks, extrapolation of this trend predicts that within 5 years, AI systems will be capable of automating many software tasks that currently take humans a month.” Established By v4 (2026-07-10) the paper's title had changed from “Measuring AI Ability to Complete Long Tasks” to “Measuring AI Ability to Complete Long Software Tasks”, with the abstract text unchanged. Frontier A brief written from a different version therefore describes a materially different claim, and this one names its version for that reason.

Established The human baselines underneath the time-horizon result are unusually good, and one of them contains a reversal that is almost never quoted. HCAST (arXiv:2503.17354) supplies 189 tasks across ML engineering, cybersecurity, software engineering and general reasoning, with 563 human baselines totalling over 1500 hours, humans working “under identical conditions as AI agents”; agents “succeed 70-80% of the time on tasks that take humans less than one hour, and less than 20% of the time on tasks that take humans more than 4 hours.” Established RE-Bench (arXiv:2411.15114) ran 71 eight-hour attempts by 61 distinct human experts across seven open-ended ML research environments, 24% of which matched or beat the reference solution. Established The quoted result is that “the best AI agents achieve a score 4x higher than human experts when both are given a total time budget of 2 hours per environment.” Established The unquoted result, from the same abstract, is that “humans currently display better returns to increasing time budgets, narrowly exceeding the top AI agent scores given an 8-hour budget, and achieving 2x the score of the top AI agent when both are given 32 total hours.” Frontier Read together, the honest summary is that current agents win short contests and lose long ones, and that the trend line everyone extrapolates is a trend in how fast the crossover point is moving.

Established A coherent body of negative results says the failures are kind-limited rather than difficulty-limited, and it is not a fringe literature. Dziri et al., arXiv:2305.18654, find across multi-digit multiplication, logic grid puzzles and dynamic programming that transformers “solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills.” Established McCoy et al., arXiv:2309.13638, report the cleanest single number in the sceptical literature: GPT-4's accuracy on decoding a simple cipher is “51% when the output is a high-probability word sequence but only 13% when it is low-probability” — on a deterministic task where output frequency should be irrelevant. Established Mirzadeh et al., arXiv:2410.05229, show that changing only the numeric values in grade-school word problems degrades every model tested, and that “adding a single clause that seems relevant to the question causes significant performance drops (up to 65%) across all state-of-the-art models.” Frontier Against that, Sinha et al., arXiv:2509.09677, argue that much long-horizon failure is execution failure rather than reasoning failure: marginal gains in single-step accuracy compound into exponential gains in task length, and models “become more likely to make mistakes when the context contains their errors from prior turns” — a self-conditioning effect that scale alone does not fix but explicit reasoning does mitigate. Frontier These two readings are not yet decidable against each other, and saying so is more useful than choosing.

Established Capability forecasting has a track record, and the track record is mostly a record of forecasts moving. Grace et al., arXiv:2401.02843, surveyed 2,778 researchers who had published at top AI venues. Established Aggregate forecasts: a 10% chance of machines outperforming humans on all tasks by 2027 and a 50% chance by 2047; full occupational automation at 10% by 2037 and 50% by 2116. Established The 2047 median is 13 years earlier than the same survey's 2022 result — a thirteen-year revision in two years, which is a fact about the instrument at least as much as about the world. Established On risk, 68.3% of respondents thought good outcomes more likely than bad, while between 38% and 51% gave at least a 10% chance to outcomes “as bad as human extinction.” Frontier No published calibration study for expert AI timelines was located for this brief, so the survey should be read as an aggregated opinion with a documented drift rate, not as a forecast with a scoring history.

3 · Frontier questions

Frontier The most serious current attempt to define AGI operationally grounds it in human psychometrics and reports a profile rather than a threshold. Hendrycks et al., arXiv:2510.18212, build on Cattell–Horn–Carroll theory, decompose general intelligence into ten core cognitive domains, and report composite scores of 27% for GPT-4 and 57% for GPT-5. Frontier The qualitative claim matters more than the composite: modern systems show a “highly jagged cognitive profile” with “critical deficits in foundational cognitive machinery, particularly long-term memory storage.” Speculative If the jaggedness is stable, “AGI achieved or not achieved” is a category error and the object of measurement is a vector; if the memory deficit vanishes when a retrieval store is attached, the jaggedness was in the harness, and nothing this brief could reach settles which.

Frontier The best teaching case for protocol discipline is an exchange that happened inside six months. Shojaee et al., arXiv:2506.06941, report three regimes on controllable puzzles — standard models ahead at low complexity, reasoning models ahead at medium, both in “complete collapse” at high — with reasoning effort declining at the top despite remaining token budget. Frontier Lawsen's comment, arXiv:2506.09250, answers that runs risked exceeding output token limits with models saying so, that the automated grader “fails to distinguish between reasoning failures and practical constraints,” and that the River Crossing benchmarks “include mathematically impossible instances for N > 5 due to insufficient boat capacity, yet models are scored as failures for not solving these unsolvable problems.” Frontier The sceptical literature has its own protocol failures, and a brief that audits only the enthusiasts is doing half the job.

Frontier Inference-time compute is a scaling axis in its own right, and it makes most reported percentages uninterpretable. Brown et al., arXiv:2407.21787, show coverage — the fraction of problems solved by any sample — scaling log-linearly with sample count across four orders of magnitude. Established On SWE-bench Lite, DeepSeek-Coder-V2-Instruct rises from 15.9% with one sample to 56% with 250, beating the then single-sample state of the art of 43%. Established Where no automatic verifier exists, majority voting and reward models “plateau beyond several hundred samples.” Frontier Repeated sampling converts compute into capability only where checking is cheap, which is a strong argument that agentic progress is concentrated in verifiable domains.

Frontier Two credible contamination results point in opposite directions, and the tie stands. Li and Flanigan, arXiv:2312.16337: “for classification tasks with no possibility of task contamination, LLMs rarely demonstrate statistically significant improvements over simple majority baselines, in both zero and few-shot settings.” Frontier Zhang et al., arXiv:2405.00332, built GSM1k as a matched clean replacement for GSM8k, found drops of up to 8% and family-specific overfitting, but also that “many models, especially those on the frontier, show minimal signs of overfitting.” Frontier Both are well constructed; declaring the tie is the terminal position here.

Frontier The single most important unresolved factual question in this subject is whether the ARC-AGI public training set was in o3's pre-training. It is asserted constantly in commentary and was established by nothing this brief's research could read. Frontier It is stated here as an open question in both directions, because it decides whether the o3 ARC result measures generalisation or preparation — which is exactly Chollet's own distinction, applied to the benchmark he built to test it. Speculative A public corpus-membership audit would settle it and nobody has published one.

Frontier The unsaturated benchmarks are where the disagreement now lives, and they move faster than the papers about them. FrontierMath (arXiv:2411.04872) reported that at launch “current state-of-the-art AI models solve under 2%” of its new, unpublished, automatically verified problems; who funded it and who had access to those problems before evaluation is not established by anything read for this brief. Frontier Humanity's Last Exam (arXiv:2501.14249) states its rationale flatly — “LLMs now achieve over 90% accuracy on popular benchmarks like MMLU” — and reported low accuracy and low calibration at launch across 2,500 expert-written questions. Established It reached its eleventh version by 2026-07-28, eighteen months after the first.

Frontier The nearest positive datum on recursive self-improvement is in-house and unreplicated. AlphaEvolve (arXiv:2506.13131), an LLM-driven evolutionary coding agent, reports a production data-centre scheduling heuristic, a circuit simplification, faster training of its own underlying model, and 4×4 complex-valued matrix multiplication in 48 scalar multiplications — “the first improvement, after 56 years, over Strassen's algorithm in this setting.” Frontier The self-acceleration claim is the one that matters here, it is the one with no external replication, and RE-Bench still shows human experts ahead at long time budgets.

Frontier Whether there is a general factor across machine systems at all is now an empirical programme rather than a slogan. Ilić and Gignac, arXiv:2310.11616, tested 591 language models on 12 instruments and report “strong empirical evidence for a positive manifold and a general factor of ability,” with parameter count positively but diminishingly related to it. Frontier Their own title carries the hedge that decides the question — “Indications of artificial general intelligence or achievement?” — because a positive manifold across models trained on overlapping web corpora may index what was in the data. Speculative The controls that would separate the two — disjoint corpora, tasks built after all training cutoffs, removal of the scale covariate — are absent, and no measurement-invariance test of a machine general factor across model families surfaced at all.

4 · Technological bottlenecks

Established The workback target is not a system but an instrument: a measurement of general intelligence that, applied unchanged to a model released after the measurement was designed, yields a number predicting that model's performance on tasks nobody has thought of yet. Every instrument surveyed here fails that test. Established They are saturated (MMLU), gamed by sampling volume (any pass@k figure), plausibly contaminated, or so new that no post-hoc model has been run on them. Frontier The chain below is ordered; the binding link is the fourth, and it is binding because nobody has attempted it.

Frontier L1, item-level psychometric infrastructure. Per-item difficulty and discrimination parameters, fitted across many models and published as a reusable instrument. Established Partially present: tinyBenchmarks and Anchor Points both show that on the order of 100 well-chosen items reproduce MMLU rankings, which is the same underlying fact, but neither releases calibrated item parameters. Frontier L2, a generative item model. A procedure that emits new items at specified difficulty from the L1 parameters, so the test can be regenerated after every model release. Established The functional-benchmark method exists for mathematics, where re-instantiating the same problem exposes reasoning gaps of 58.35% to 80.31% in models that score well on the static version; it has not been generalised beyond arithmetic templating to semantic transformation.

Frontier L3, contamination-proof administration. Items that provably did not exist before the model's training cutoff, administered once. Established FrontierMath's private-set protocol and ARC-AGI's semi-private and private splits are working instances; both are expensive and both carry unresolved independence questions. Frontier L4 is the binding link: construct validity. Nobody has demonstrated that a latent ability estimated on one instrument predicts out-of-distribution performance on tasks from a genuinely different generator. Frontier Until that exists, every machine-intelligence score is a description of a test rather than a measurement of a system, and Raji et al.'s construct-validity objection to “general” benchmarks stands unanswered.

Frontier L5, invariance under scaffolding. If the variance in agent results attributable to the harness exceeds the variance attributable to the model, the object being measured is not the model. Established Kapoor et al. document exactly this confound — benchmarks that score accuracy while ignoring cost, conflated developer needs, benchmarks with no holdout set at all, and pervasive irreproducibility — but do not decompose the variance. Frontier L6, cross-population human anchoring. HCAST-grade timed baselining exists for software and for nothing else. Frontier L7, the survival test: apply the finished instrument, unchanged, to three models released after its publication and report rank-order and calibration. Established Nothing in this literature has passed L7, and that is the field's central negative finding rather than an oversight.

5 · Research dependencies

Established This brief depends first on measurement, and the dependency is unusually literal. Almost every claim about general intelligence in circulation is a claim about a benchmark result, so the reliability of the benchmarks bounds the reliability of the subject. Established That work — item response theory for machines, contamination auditing, reliability metrics such as pass^k, error bars on evaluations, and the reproducibility of harnesses — belongs to Intelligence Measurement and is treated there. Frontier Until it lands, this subject cannot distinguish a capability gain from a measurement artifact, and the history in section 14 shows it repeatedly failing to.

Frontier It depends second on human baselining outside software. The time-horizon programme is rigorous precisely because HCAST and RE-Bench timed real experts under matched conditions; no comparably rigorous human-baselined time-horizon measurement exists for wet-lab work, litigation, clinical practice or skilled trades. Speculative The absence claim is stated as an absence found rather than a proof: no counterexample surfaced in the research for this brief. Frontier Until it does, extrapolating a software doubling time to the economy is an inference across a boundary nobody has measured.

Frontier It depends third on interpretability, and fourth on results this brief cannot reach. A mechanistic account of whether a model has internal structure corresponding to a task's solution, rather than to a surface correlate of it, would decide the Faith-and-Fate and Embers results directly instead of by behavioural inference. Speculative And a corpus-membership audit for benchmark training sets — a database question, not a research question — is a dependency that could be discharged by any lab willing to publish it, which makes it the one item on this list blocked by incentive rather than by capability. Frontier A fifth dependency is unglamorous and probably decisive: sustained funding for private evaluation splits, since every contamination-proof instrument in this subject is expensive to build, cheap to burn, and worthless the moment its items leak.

6 · Required experiments

Frontier The decisive experiment is a pre-registered out-of-distribution prediction, and it costs almost nothing except institutional willingness. Fit the latent ability of N models on instrument A; publish point predictions for their scores on instrument B before instrument B exists; have instrument B built by a disjoint team from a disjoint task generator; report calibration. Frontier No such pre-registered prediction from a machine psychometric ability estimate surfaced in the research for this brief, and its absence is the reason the field can argue indefinitely.

Frontier The cheapest high-value experiment is the corpus audit. Publish, for one frontier model, whether the ARC-AGI public training set appears in pre-training, with the matching methodology open. Established Naive string-matching decontamination is known to be defeated by paraphrase and translation, so the audit must use the rephrasing-robust methods the contamination literature already developed. Frontier A negative result would make the o3 ARC numbers a generalisation result; a positive one would make them a preparation result; the current state, where commentary asserts both, is the worst of the three.

Frontier Then the scaffold-variance decomposition. Run the same models under five independently built agent scaffolds at fixed compute and report the variance component attributable to the scaffold. Frontier If scaffold variance dominates model variance, most published agent comparisons are measuring engineering teams. Frontier Then the reliability curve. METR's horizon is defined at 50% success; no published 95%-reliability time-horizon curve surfaced, and a system that finishes month-long tasks half the time is economically a different object from one that finishes them nineteen times in twenty. Frontier Then the compute sweep. ARC-AGI-2 has not yet been measured across two orders of magnitude of inference compute; the discontinuity hypothesis requires flatness in compute, and the benchmark is too young for anyone to claim it. Speculative Finally, extend HCAST-grade timed human baselining to one non-software domain — a single well-run wet-lab, clinical or legal replication, with experts timed under conditions matched to the agent's, would test the external-validity assumption on which the whole time-horizon extrapolation rests, and would cost a small fraction of what one frontier training run costs. Frontier Note what all five of these have in common: none requires a new model, a new architecture or a new idea. They require somebody to run the measurement and publish it before the result is known.

7 · Engineering requirements

Established The engineering constraint that most changes the picture is inference cost, and it is routinely omitted from the numbers. The ARC-AGI-2 paper's own figures make the point without commentary: the same system, at two compute settings, cost $200 and $20,000 per task and differed by twelve percentage points. Frontier A capability that exists at $20,000 per task exists in a laboratory sense and not in an economic one, and a benchmark table without a cost column cannot distinguish the two.

Frontier Second, verifiability is an engineering property of the deployment domain, not of the model. Repeated sampling converts compute into capability where an automatic verifier exists and plateaus where it does not. Frontier That makes formal verification, test generation and simulation environments load-bearing infrastructure for general capability rather than adjacent conveniences, and it predicts that agentic progress will stay concentrated in code, mathematics and games until cheap verifiers exist elsewhere.

Frontier Third, memory. The psychometric profile that names long-term memory storage as a critical deficit is naming an architectural gap, and retrieval augmentation is the current patch rather than a solution to it. Speculative Whether persistent, consolidating memory is a model property or a systems property is unsettled, and it is the question whose answer determines whether the jagged profile is a fact about transformers or about harnesses. Frontier Fourth, and against the grain of most roadmaps, a hardware vendor's position paper (arXiv:2506.02153) argues that most agentic invocations are “specialized tasks repetitively and with little variation” and that small models are therefore “sufficiently powerful, inherently more suitable, and necessarily more economical” for them; the framing serves the author's commercial position and the argument is still worth answering. Speculative If it is right, the deployed shape of machine intelligence is a swarm of narrow models under an orchestrator, general capability is a property of the system rather than of any model in it, and the entire practice of scoring individual models as candidates for general intelligence is measuring the wrong object.

8 · Adjacent technologies

Frontier The orchestration route to general intelligence is currently losing, and that is under-reported. The engineering-integration hypothesis holds that the components — perception, retrieval, planning, execution, tool use — are all adequate and that general intelligence is what good enough orchestration produces. Frontier The multi-agent literature reviewed in Multi-Agent Intelligence Systems reports that orchestration gains on popular benchmarks are often minimal, and no compute-matched controlled comparison of multi-agent against single-agent systems surfaced in this research at all. Frontier Until one exists, claims that multi-agent scaffolding produces qualitatively new capability are claims about uncontrolled comparisons.

Frontier Autonomous scientific discovery is the adjacent field where a general-capability claim would be hardest to fake, and it is treated in Artificial Scientists. A system that produces a genuinely novel, independently verified result nobody prompted it toward would be evidence no benchmark can supply. Frontier Artificial Creativity supplies the complementary distributional finding that recurs across this cluster: models beat the human average and lose the human right tail. Speculative If that pattern holds for reasoning as well as for idea generation, it is a precise description of what current systems are — competent everywhere, exceptional nowhere — and it is not the shape most AGI narratives assume.

Frontier Two further adjacencies matter and one of them should be excised. Cognitive Architectures holds the older symbolic and hybrid programmes whose vocabulary the psychometric definitions have quietly re-imported. Speculative Machine Consciousness holds a question that is orthogonal to this one and should be kept out of it: none of the benchmarks in this brief measure anything a philosophical zombie could not do. Handwave That observation cuts both ways, and the sharper edge is the uncomfortable one — if no current instrument can distinguish a general intelligence from a very good zombie, the instruments may not be measuring the thing people actually care about.

9 · Institutional requirements

Established The intergovernmental machinery exists and its structure, at least, is a matter of record. The International AI Safety Report (arXiv:2501.17805), chaired by Bengio, was mandated by the states at the Bletchley summit; 30 nations plus the UN, OECD and EU nominated Expert Advisory Panel members, around 100 experts contributed, and those experts “collectively had full discretion over the report's content.” Frontier That is a governance fact, not an evaluation result, and the distinction is worth holding: this brief could not retrieve any national institute's evaluation methodology or numbers, and states none.

Established The field's foundational capability claims were, and largely still are, published by interested parties. The “Sparks of AGI” paper studied “an early development version of GPT-4” — not the shipped model — was written entirely by researchers at Microsoft Research while Microsoft was OpenAI's principal investor, and used no held-out protocol and no baseline. Established The bar-exam percentile originated in OpenAI's own technical report. Frontier Neither is fraud and both are normal industrial practice; the institutional problem is that the field's shared vocabulary was set by documents whose authors had a position in the outcome, and the corrections arrived years later in venues with a fraction of the readership.

Frontier What is missing is an instrument custodian with the standing of a metrology institute. The functions are identifiable: hold private evaluation splits, run administrations under declared protocols, publish item parameters and confidence intervals, and refuse to certify a number without its compute cost and attempt count. Speculative No such body exists with the independence, funding and access required, and the funding and problem-access arrangements of existing private benchmarks are in several cases not publicly established. Frontier The governance of who is permitted to declare that a threshold has been crossed belongs to AI Governance, and it is a live question precisely because the declaration would move capital and law before any instrument could check it. Speculative The nearest working analogy is metrology rather than peer review: national standards institutes do not evaluate every product, they hold the reference and certify the procedure, and no equivalent reference exists for machine capability. Frontier Building one is an institutional problem with a known shape and no owner, which is a different and more tractable situation than a scientific problem with no solution.

10 · Ethical & societal considerations

Established Overstated capability claims are not a victimless publishing habit; they set policy about people. The bar-exam figure entered public argument as evidence that a system had reached the top decile of a licensed profession, and the correction — roughly 62nd percentile against first-time takers, roughly 48th against those who actually passed, and roughly 15th on essays against that same group — arrived in a law-and-AI journal two years later. Established Claims about professional displacement, procurement decisions and educational policy were made in the interval. Frontier The ethical failure is specific and repeatable: a percentile was quoted with no reference population named, and no gate in the publication chain required one.

Frontier The same defect runs through labour forecasting. The time-horizon extrapolation most widely cited as evidence that month-long human work will be automated is conditional in its own text, scoped to software by its own title, defined at 50% reliability, and built on a doubling trend its authors flag as possibly unstable. Frontier Every one of those four qualifications is load-bearing and all four are routinely dropped. Speculative A worker or a legislator acting on the unqualified version is acting on a claim the source does not make.

Established On catastrophic risk the survey evidence is genuinely split and should be reported split. Among 2,778 published AI researchers, 68.3% thought good outcomes more likely than bad, while between 38% and 51% assigned at least a 10% chance to outcomes as bad as human extinction. Frontier Those two findings are held by overlapping populations and are not contradictory; a field can be net optimistic and hold a substantial tail. Speculative Reporting either number alone — and both are reported alone, by opposite constituencies — misrepresents a survey that took care to ask both questions. Frontier The deeper ethical point is that a field with no instrument capable of surviving contact with the next model is being asked to advise on decisions whose horizon is decades, and the honest posture under that constraint is to state the uncertainty at full width rather than to supply the confident number the decision would prefer.

11 · Civilizational implications

Speculative If the continuity position is right, the civilizational question is about rate and distribution rather than about a threshold. That position requires three things to hold: that loss power-laws continue across the next two to three orders of magnitude of compute, that downstream capability remains a smooth monotone function of loss, and that the tasks still failing are difficulty-limited rather than kind-limited. Frontier The strongest evidence for it is that mere repeated sampling moved a coding benchmark from 15.9% to 56%, implying the capability was latent and only elicitation-limited. Frontier The strongest evidence against is that failures track output probability and surface perturbation rather than logical difficulty, which is a signature of kind-limitation.

Speculative If the jagged-profile reading is right, the civilizational consequence is stranger and less discussed. A world with systems that are superhuman at retrieval and synthesis and sub-human at forming durable memory does not get a moment when general intelligence arrives; it gets a long period in which institutions must decide, task by task, which side of the jag they are on. Speculative Planning for a threshold in that world produces exactly the wrong institutional design.

Handwave And if the recursive-self-improvement framing is right, everything above is a distraction. On that view the only threshold with discontinuous consequences is a system that improves its successor faster than humans can, and the AGI question as usually posed is mis-specified. Frontier The nearest measurement instrument shows human experts still ahead at long time budgets; the nearest positive datum is one lab's report that its own agent accelerated the training of its own underlying model, unreplicated and in-house. Handwave That is a thin evidential base for the most consequential hypothesis in the subject, and the thinness is the point: the scenario with the largest stakes is the one with the least measurement attached to it.

12 · Timelines

These horizons track measurement rather than capability, because measurement is what this subject is currently short of. Capability dates are given as the field's own aggregates, with the aggregate's drift rate attached.

  • 10 yr: Frontier An instrument passing the survival test — applied unchanged to three subsequently released models with rank-order and calibration holding — is achievable this decade and requires no new science, only a custodian and a pre-registration habit. Frontier Non-software human baselining at HCAST grade in at least one domain is comparably tractable. Speculative Expect the definitional dispute to remain unresolved, because nothing on this list adjudicates between skill-acquisition efficiency and depth-times-breadth.
  • 25 yr: Frontier This is the window containing the surveyed 50% aggregate for machines outperforming humans at all tasks, which stood at 2047 in the 2024 survey and had moved thirteen years earlier in two years. Speculative Treat the date as a moving aggregate rather than a forecast; the informative quantity is the drift, not the year. Speculative If the jagged-profile reading holds, the question that gets answered in this window is which faculties remain jagged after memory architectures mature, not whether a threshold was crossed.
  • 50 yr: Speculative The surveyed aggregate puts full occupational automation at 50% by 2116, which places most of this window before the median rather than after it. Speculative A plausible outcome is that the term is retired: measurement moves to capability vectors, procurement and regulation attach to specific reliability thresholds in specific domains, and “AGI” survives as a marketing category rather than a technical one. Handwave The alternative — a documented, replicated recursive-self-improvement result — would make every date on this list irrelevant, and there is currently no instrument that would detect it early.
  • 100 / 250+ yr: Handwave At this range the honest content is conditional structure rather than dates. Handwave If general intelligence turns out not to be a natural kind, the historical question becomes why a folk concept organised a century of research funding, and the answer will be about institutions rather than about cognition. Handwave If it does turn out to be one, the interesting artefact from this era will be that the field spent its first decades unable to build a ruler that survived contact with the next model.

13 · Technology tree & dependencies

  • Depends on This brief depends on Intelligence Measurement more completely than most dependencies on this map, because nearly every proposition about general intelligence in circulation is a proposition about a benchmark result. Item response theory fitted across models, contamination auditing robust to paraphrase, reliability metrics that report the probability of succeeding on all of k trials rather than any of them, confidence intervals on evaluations, and reproducible harnesses are all prerequisites here rather than refinements. The dependency is asymmetric and worth stating plainly: better measurement could show that the capability gains of the last five years are smaller than reported, and no amount of capability work can show that the measurements were sound. The second dependency is on human baselining outside software, which is a logistics and funding problem rather than a scientific one, and which nobody has funded. The third is a corpus-membership audit for benchmark training data, which is a database query that any frontier lab could publish and none has.
  • Enables What this brief enables is mostly discipline exported to its neighbours. A capability-vector framing rather than a threshold framing changes what Artificial Scientists has to demonstrate, what Multi-Agent Intelligence Systems has to control for, and what AI Governance can sensibly attach a rule to: reliability thresholds in named domains rather than a declared arrival. It supplies the standing correction that any brief quoting a benchmark percentage must carry the protocol with it, which is a habit rather than a result but is the habit this cluster most needs. And it supplies a negative enabling condition that is genuinely useful: because none of the instruments here measure anything a system without inner experience could not do, work on machine consciousness cannot borrow a capability result as evidence, and does not have to. The most concrete thing it enables, though, is a refusal. A brief, a procurement document or a regulation that quotes a benchmark percentage without its evaluation split, its attempt count and its inference cost is now quoting a number this map treats as unfinished, and the o3 case gives that refusal a worked example rather than a principle.
  • Adjacent Intelligence Measurement, Multi-Agent Intelligence Systems, Artificial Scientists, Artificial Creativity, Cognitive Architectures, Machine Consciousness, AI Governance, Collective Intelligence. The measurement brief is the dependency; the rest are neighbours that share this brief's failure modes. Artificial Scientists shares the novelty-verification problem, Artificial Creativity shares the beats-the-average-loses-the-tail distribution, Multi-Agent Intelligence Systems shares the missing controlled comparison, and AI Governance inherits the whole problem of who is entitled to declare that something has arrived.

14 · Common misconceptions & speculative claims

Established “o3 solved ARC-AGI.” This is the most consequential error in the subject and the correction is documented inside a single paper. Established The ARC-AGI-2 paper's body reports the December 2024 o3 preview at 76% (low compute, estimated $200 per task) and 88% (high compute, estimated $20,000 per task) on the ARC-AGI-1 semi-private evaluation set. Established The same paper's Table 1 reports the released o3 (Medium) at 53.0% on that same set, and 3.0% on ARC-AGI-2. The ARC Prize target, separately, is 85% on the private set under a compute cap. Frontier So the demonstrated model is not the shipped model, the split quoted in commentary is often not the split the target refers to, and the two demonstration figures differ in cost by two orders of magnitude. Established Note also that the widely circulated “75.7% / 87.5%” pair does not appear in that paper at all; anyone quoting the more precise pair is sourcing it from an announcement this brief could not retrieve.

Frontier “The ARC training set was in o3's pre-training” — and also its denial. Both are asserted routinely and neither was established by anything this brief's research could read. Frontier It is stated here as an open question because it is the question that decides what the number means: on Chollet's own definition, a system prepared on a task family and a system generalising to it are not the same kind of thing, and this is his benchmark, built to test exactly that distinction. Speculative Anyone who wants to end the argument can publish a corpus-membership audit using paraphrase-robust matching, and the fact that nobody has is a fact about incentives rather than about difficulty.

Established “GPT-4 passed the bar exam in the top 10%.” The 90th-percentile figure is OpenAI's own, from its GPT-4 technical report. Established Martínez, in Artificial Intelligence and Law, recomputes it against named reference populations: approximately the 62nd percentile against first-time takers overall, approximately the 48th against licensed or license-pending attorneys, and on the essay components approximately the 42nd against first-time takers and approximately the 15th against those who passed; the abstract states that OpenAI's estimates “are overinflated.” The enthusiast error is not that the number is wrong but that a percentile was quoted with no reference population named. Established The correction has not propagated: when the English Wikipedia article on stochastic parrots was fetched during this research, it still listed “GPT-4 scored in the >90th percentile on the Uniform Bar Examination” among the published counterarguments to Bender et al. Frontier A two-year-old peer-reviewed correction has not reached the entry where the refuted figure is doing argumentative work.

Established “AI will automate month-long human tasks by 2030.” METR's own sentence is conditional and scoped: if the results generalise to real-world software tasks, extrapolation predicts automation of many software tasks that currently take humans a month within five years. Established Software; 50% success; a seven-month doubling the authors themselves flag as possibly accelerated in 2024, which makes the trend line less reliable to extrapolate rather than more. Established And the paper's title changed between versions: “Measuring AI Ability to Complete Long Tasks” at v2, “Measuring AI Ability to Complete Long Software Tasks” at v4, with an identical abstract. Frontier A brief written from a different version describes a different claim; this one was written from v2 and says so.

Established “Emergent abilities prove that scale produces qualitative jumps” — and the deflationary overreach that answers it. The jumps largely track metric discontinuity, in-context learning and model memory. Established But Schaeffer et al. show that alleged emergent abilities evaporate under continuous metrics, which is not a demonstration that no capability ever appears discontinuously. Frontier Both sides claim more than the papers support.

Established “Chinchilla showed we should train smaller models” and “Kaplan showed we should train bigger ones.” Kaplan's result is about sample-efficiency under a fixed compute budget; Chinchilla's is that parameters and tokens should scale equally. Established Neither says anything about inference cost, which is what determines deployment economics now that repeated sampling and long reasoning traces are standard.

Established “Sparks of AGI was a study of GPT-4.” It studied an early development version the public never had, at the principal investor in the model's developer, with no held-out protocol and no baseline. Frontier It is a qualitative case-study collection cited as though it were an evaluation.

Established “AI experts now predict AGI by 2027.” The 2027 figure in Grace et al. is a 10% probability; the 50% figure is 2047. Established Quoting the 10% year as the prediction inverts the survey.

Established “Benchmark X is saturated, therefore models are superhuman at X.” MMLU's saturation is partly an artifact of its errors: an audit estimates that 6.49% of MMLU questions contain errors, rising to 57% of the analysed Virology subset. Frontier A ceiling can be the ceiling of the ruler.

Established Skeptic-side error, for symmetry: “reasoning models collapse on Tower of Hanoi, therefore they cannot reason.” A portion of the reported collapse was output-token truncation, and part of the River Crossing suite contained mathematically impossible instances that were scored as model failures. Frontier The sceptical literature has protocol failures of exactly the kind it diagnoses in the enthusiast literature.

Established “AlphaEvolve did worse than AlphaTensor: 48 multiplications versus 47.” The two numbers are not the same quantity: AlphaTensor's 47 is for 4×4 matrices over finite fields, AlphaEvolve's 48 is for 4×4 complex-valued matrices, and each paper states its own setting in its own abstract. Frontier The comparison is a unit error that enthusiastic and hostile readers make equally often.

Handwave The exotic misconception, stated in its strongest form: that “general intelligence” is a coherent target at all. Human general intelligence is a factor extracted from a population of organisms with a shared developmental architecture and a shared task ecology; neither condition holds across model families. Speculative Raji et al. argue that benchmarks “operate as stand-ins for a range of anointed common problems” and are “frequently framed as foundational milestones on the path towards flexible and generalizable AI systems” while suffering construct-validity failures that undermine the generality claim they are used to support. Frontier The strongest empirical answer is a positive manifold across 591 models whose own authors ask whether they found ability or achievement. Handwave If the folk-concept reading is right, the sensible programme is not to build a general intelligence or to measure one but to enumerate capability vectors and attach institutions to named reliability thresholds — and the argument about whether the threshold has been crossed will read like an argument about whether a whale is a fish.