1 · Concept overview
Established Every AI governance instrument now in force delegates the hard question to an evaluation science that does not yet exist. The question is is this model dangerous, and no statute answers it. The EU Artificial Intelligence Act, Regulation (EU) 2024/1689, entered into force on 1 August 2024. Established Its heaviest model-level obligations attach to general-purpose systems whose training required more than 1025 floating-point operations. Frontier That trigger is a quantity of arithmetic, not a capability, and it was chosen because arithmetic is countable and capability is not.
Established The bodies built to supply the missing measurement are recent, small and unstable. The first national AI safety institutes date from November 2023; the UK’s was renamed the AI Security Institute in February 2025 and the US body became the Center for AI Standards and Innovation in June 2025. Established Declared budgets run from GBP 100 million down to single-digit millions, and the two staff counts on the public record are roughly 23 and at least 30 people. Frontier As of April 2024 many developers had not granted pre-deployment access to their most advanced models, and no pre-deployment evaluation report from a national institute appears on the record this brief could obtain.
Speculative This brief treats AI governance as a measurement problem rather than a political one. The political theory is covered elsewhere in the corpus — see Institutional Design, Scientific Advisory Institutions and Existential Risk Governance — and is deliberately kept light here. Frontier What follows is the technical-institutional question: what would a competent authority have to be able to measure, what does it currently measure, and what would have to be true for the gap to close. Handwave The exotic end of the subject is the possibility that no such measurement exists, in which case the entire apparatus is a bet on a proposition nobody has tested.
2 · Current scientific position
Established The instrument, and the five dates it can honestly be said to have. The EU Artificial Intelligence Act is Regulation (EU) 2024/1689. The European Parliament voted in March 2024; the Council adopted on 21 May 2024; the Regulation itself bears the date 13 June 2024; it was published in the Official Journal on 12 July 2024; and it entered into force on 1 August 2024. The date almost always reported is 21 May 2024, and the date that belongs in a legal citation is 13 June 2024, because that is the date the instrument carries. Frontier This brief takes the 13 June date from the title string of the International Legal Materials introductory note by Nathalie A. Smuha, verified as a bibliographic record but not read here.
Established The staged application, in the only form this brief can source. The Act phases in over 6 months for bans on unacceptable-risk systems, 9 months for codes of practice, 12 months for general-purpose AI, 36 months for some high-risk obligations and 24 months for everything else, counted from entry into force. Established One absolute date is independently attested: the general-purpose AI model provisions apply from 2 August 2025. Frontier This brief prints no other calendar date for the Act’s staging, and the omission is deliberate. The allocation is set by Article 113, which this brief’s research pass could not obtain — the Official Journal text, the Commission’s own timeline page and the consolidated article-level text were all unreachable. Arithmetic from 1 August 2024 produces plausible dates; it does not tell you which Chapter sits on which of them, because Article 113’s carve-outs are article-by-article. Speculative A brief that guessed here would be doing precisely what it criticises regulators for: substituting a number that can be computed for a fact that has to be checked.
Established The architecture: four tiers, plus a fifth regime that does not fit them. Uses are sorted into unacceptable risk (prohibited), high risk (security, transparency and quality obligations), limited risk (transparency obligations) and minimal risk (no obligations at all). General-purpose AI sits under a separate framework. Frontier The structural fact that matters is that the general-purpose regime is orthogonal to the risk pyramid: it was added after the release of general-purpose chat systems, and it regulates a model rather than a use. Frontier Those are different regulatory objects with different measurement needs: a use can be described, enumerated and inspected in situ, while a model is a set of weights whose behaviour depends on what is bolted to it after certification. Established Within the general-purpose regime, models are treated as posing systemic risk where training required more than 1025 floating-point operations. Penalties run to EUR 35,000,000 or 7 per cent of worldwide annual turnover for prohibited practices, EUR 15,000,000 or 3 per cent for other operator obligations, and EUR 7,500,000 or 1 per cent for supplying misleading information.
Established The machinery that turns obligations into engineering, and the date it is due. The Act’s high-risk obligations become operable through harmonised standards drafted by CEN/CENELEC Joint Technical Committee 21. Established Once those standards are published in the Official Journal, compliant products are presumed to conform with the Regulation. Frontier That presumption is the whole operational content of the high-risk regime: until the standards land, the obligations are abstractions no engineer can build against and no assessor can audit against. Speculative This brief cannot date the standards. No source it could reach gives a JTC 21 delivery deadline, a Commission deadline, a delay, or the content of the amendment package proposed in late 2025. Handwave A delay is very widely discussed and is not asserted here. The gap between an obligation’s date and the standard that operationalises it is the exact thing this brief exists to examine, and it deserves a citation rather than an impression.
Established The evaluators: roster, founding dates, money. National institutes were founded in the UK and US in November 2023, Japan in February 2024, Singapore in May 2024 (renamed from a Digital Trust Centre founded in June 2022), Canada and South Korea in November 2024, France on 31 January 2025 as INESIA, India on 30 January 2025, and Australia on 25 November 2025. The EU AI Office dates from May 2024. Kenya is a network member and the only African state in it, with no institute details announced. Established Declared funding, as recorded: GBP 100 million initial for the UK body in its earlier form as the Frontier AI Taskforce, USD 10 million for the US body in March 2024, CAD 50 million over five years for Canada, SGD 10 million per year for Singapore, and INR 20 crore for India. Frontier A figure for South Korea appears in the same table in units that are almost certainly wrong by several orders of magnitude, and this brief does not print it. Staff: Japan roughly 23 people; South Korea at least 30. Frontier That table is the most important quantitative fact in the subject. The declared public evaluation capacity of the entire world, summed, is smaller than the compute bill for a single frontier training run.
Established The access problem, stated by the institutes’ own chronicle. As of April 2024, many AI companies had not shared pre-deployment access to their most advanced models for evaluation, and no published pre-deployment evaluation reports are recorded. Established The principal published methodology attributable to a national institute is Inspect, open-sourced by the UK in May 2024, which evaluates model capabilities such as reasoning and degree of autonomy. Established And black-box access — query-only, which is what an evaluator without weights has — is formally insufficient. Casper and twenty co-authors distinguish black-box, white-box (weights and gradients) and outside-the-box (development documentation) access, and show that white- and outside-the-box access allow substantially more scrutiny than black-box alone; they add that transparency about the access and methods used is necessary to interpret an audit result at all. Frontier Put those together and you have this brief’s thesis in one line: the public evaluators hold the access level the audit literature says is inadequate, and nobody has published what that costs in detection power.
Established What the field says should be measured. The specification document is Shevlane and twenty-two co-authors from DeepMind, OpenAI, Anthropic, Oxford, Cambridge and Bengio’s group, which splits the problem into dangerous capability evaluations (what can the model do — offensive cyber capability, strong manipulation) and alignment evaluations (would it apply them). Established The only worked example of such a battery at frontier scale on the public record is Phuong and twenty-seven co-authors at Google DeepMind, covering persuasion and deception, cyber-security, self-proliferation and self-reasoning, run on Gemini 1.0, which found no strong dangerous capabilities but preliminary warning indicators. Frontier The paper’s own stated aim is to help advance a rigorous science of dangerous capability evaluation in preparation for future models. Frontier That phrase is the admission that matters and it is made by the people best placed to know. Established The institutional counterpart is Anderljung and twenty-three co-authors: three pillars of standard-setting, registration and reporting, and compliance enforcement, with safety standards covering pre-deployment risk assessment, external scrutiny, informed deployment decisions and post-deployment monitoring.
Established Why compute became the lever. Sastry, Heim and seventeen co-authors give the reason in one sentence: computing power is detectable, excludable and quantifiable, and is produced through an extremely concentrated supply chain. Their three applications are enhancing regulatory visibility, steering resources toward good outcomes, and restricting dangerous development. Established The same authors warn that implementation readiness varies widely and that badly designed compute governance risks privacy violations, economic harm and power centralisation. Frontier Compute is not a proxy for danger; it is a proxy for the kind of activity that has historically preceded capability jumps, which is a weaker claim, and it remains the strongest observable anyone has proposed that is both cheap to measure and hard to hide.
Established The international layer exists, in two shapes. The first standing assessment body is the International AI Safety Report chaired by Yoshua Bengio, mandated by the nations attending the Bletchley summit: thirty nations plus the UN, OECD and EU each nominated a representative to its Expert Advisory Panel, roughly 100 experts contributed, and those experts held full discretion over the content. Frontier That is an IPCC-shaped institution for AI, and it is the strongest existing answer to the question of whether an authority could ever know enough. Established The second shape is treaty-like: the Council of Europe Framework Convention on Artificial Intelligence and Human Rights, Democracy and the Rule of Law, adopted 17 May 2024 and opened for signature 5 September 2024, with first signatories including Andorra, Georgia, Iceland, Norway, Moldova, San Marino, the United Kingdom, Israel, the United States and the European Union. Speculative Its core is procedural and rights-based and it creates no measurement capacity whatever, which is worth saying plainly, because it is routinely described as the first binding international AI treaty and read as though that settled something.
Established The rest of the landscape, dated. The United States issued an executive order on AI safety and security in October 2023 requiring developers to share safety-testing results with government before public release, and a further order on 12 December 2025 directing federal agencies to develop a unified national policy and to evaluate state laws for conflicts. Established China has moved by instrument rather than by statute: algorithm recommendation provisions effective March 2022, deep synthesis provisions effective January 2023, generative AI measures from 15 August 2023 described as the first comprehensive national framework, ethics review measures in October 2023, and labelling measures for AI-generated synthetic content in 2025. Established The G7 Hiroshima process produced eleven guiding principles and a voluntary code of conduct on 30 October 2023, and the summit sequence runs Bletchley November 2023, Seoul 2024, Paris 2025, New Delhi 2026. Frontier Those jurisdictions regulate three different objects — the use, the model and the output — so an evaluator seeking to serve all three must produce evidence in three currencies that do not convert.
Frontier The subject the evaluators are missing has a name, and the name is metrology. Everything above treats an evaluation as a policy instrument; the more useful framing is that an evaluation is a measurement, and measurements have a settled list of properties that must be established before the number means anything. Established A measurement needs a defined construct, a documented instrument, a reported reliability, a stated uncertainty, a traceability chain back to something stable, and a demonstrated transfer from the conditions of measurement to the conditions of use. Frontier Named that way, the AI evaluation record is not thin in one place but absent in six: published scores rarely state what construct they claim to measure, almost never carry an interval, are not traceable to any reference, and are produced under conditions that differ systematically from deployment. Established The United States metrology institute has begun issuing guidance in this register rather than the policy register, most recently an initial public draft numbered AI 800-2 in its artificial intelligence series. Frontier This brief cites that draft for its existence and its series position and prints no section numbers from it, because the text was not read here; what carries the argument is the institutional signal, which is that the body maintaining the national measurement standards has decided AI evaluation is its problem. Frontier The general machinery is carried in Metrology Infrastructure and the construct question in its pure form in Intelligence Measurement; what is added here is the observation that AI governance is the first regulatory regime in a century to be built before its metrology rather than after it.
Established Construct validity is the first of the six properties and the one most often skipped. The psychometric tradition holds that a measurement is valid only relative to an explicitly stated construct, and Jacobs and Wallach carried that apparatus into computational systems, showing that most disputes about whether a system is fair are really disputes about whether an unstated construct was operationalised correctly. Established Raji and co-authors applied the same lens to benchmarks directly, arguing that general-purpose benchmarks routinely claim a generality their task distribution cannot support. Frontier For governance the consequence is exact: the systemic-risk tier, the codes of practice and every safety case in the field rest on scores whose constructs are undeclared, so an appeal against a finding cannot be adjudicated on the merits — there is no stated thing the number was supposed to be a number of. Established Contamination is the second failure and it has been measured rather than alleged. A team that built a fresh parallel construction of a widely used grade-school mathematics benchmark found accuracy drops of up to roughly thirteen points for some model families on the new set while others were unaffected, which is the signature of memorisation rather than of capability. Frontier A regulator relying on a public benchmark is therefore relying on an instrument whose calibration decays every time the instrument is published, and nothing in any instrument described above obliges anyone to report when it has.
3 · Frontier questions
Frontier Technical AI governance now exists as a named field, and its founding document is a taxonomy of holes. Reuel and thirty-two co-authors frame the technical contribution as three functions: identifying where intervention is needed, assessing the efficacy of candidate governance actions, and designing mechanisms for enforcement, incentivisation or compliance. Frontier Thirty-three people who would rather be publishing a toolbox published an inventory of missing tools instead, which is the single best evidence for this brief’s framing.
Frontier Verification for international agreements has been itemised for the first time. Wasil, Reed, Miller and Barnett enumerate ten techniques aimed at two violation classes — unauthorised training above a compute threshold, and unauthorised data centres — grouped by intrusiveness: methods needing minimal foreign access, methods needing the suspected violator’s approval, and methods requiring advanced hardware controls. Each technique is paired with a historical analogue and with circumvention strategies, and the circumvention analysis is what makes it research rather than advocacy.
Frontier The safety case is the compliance artefact most likely to be adopted. Clymer, Gabrieli, Krueger and Larsen set out four argument families: capability absence (it cannot cause a catastrophe), control (we can stop it), trustworthiness (it will not, despite being able to), and deference to qualified advisors for exceptionally powerful systems. The form comes straight from nuclear and aviation regulation, where regulators already know how to receive and assess one, and that familiarity is its main advantage. Speculative Its weakness is that a capability-absence case is a claim about the non-existence of an elicitation over an unbounded search space.
Frontier Control protocols are the one safety claim that is auditable today. Greenblatt, Shlegeris, Sachan and Roger build a governance object rather than an alignment result: a powerful untrusted model (GPT-4), a weaker trusted model (GPT-3.5), and a limited budget of human expert review, on programming tasks where the untrusted model deliberately inserts backdoors. Trusted editing and untrusted monitoring with anti-collusion safeguards both beat the baselines. Frontier The significance is that you can measure the safety of a protocol without resolving what the model wants, converting an unanswerable question about intent into an answerable one about a deployment configuration.
Frontier The governable unit is shifting from the model to the deployed agent. The live 2025–2026 agenda is visibility into agents: identifiers, real-time monitoring and activity logging, with the privacy and power-concentration costs those measures carry. Speculative This brief flags a source problem in the open: the preprint carrying that agenda returned its substance but not its bibliographic header, so it is listed with its authorship unverified rather than cited as though the attribution had been checked. The idea does not depend on the citation. If the regulated unit is a configured, tool-using, long-running agent rather than a set of weights, almost every instrument above is aimed at the wrong object.
Frontier Adaptation is a third axis, and it may dominate. Bernardi, Mukobi, Greaves, Heim and Anderljung argue that capability control and diffusion control become increasingly impractical as developer numbers grow, and propose reducing the expected negative impact from a given level of diffusion; their interventions are to avoid, defend against and remedy, across election manipulation, cyberterrorism and loss of control to AI decision-makers. Speculative This is the strongest published argument that the measurement problem may be partly avoidable rather than solvable: you harden the target instead of certifying the weapon, and hardening does not require knowing what the weapon can do.
Frontier The expert consensus has escalated, and it says the tools are missing. Bengio, Hinton, Yao, Song, Russell, Kahneman and nineteen others, in Science, find that technical safety research lags significantly and that existing governance frameworks lack sufficient tools to address emerging autonomous systems. Speculative The interesting property of that document is not its alarm but its self-description: a consensus statement whose central empirical claim is about the inadequacy of the instruments its own signatories would have to use.
Frontier And the ruler is under suspicion from inside. Schaeffer, Miranda and Koyejo argue that emergent abilities appear due to the researcher’s choice of metric rather than fundamental changes in model behaviour with scale, and that nonlinear or discontinuous metrics produce apparent emergence where linear or continuous ones produce smooth, predictable change. Koyejo is also a co-author on the open-problems taxonomy, so the person who showed the jumps may be in the measuring instrument is co-authoring the governance agenda. Frontier If capability arrives smoothly, the jump-catching rationale for a fixed compute threshold weakens sharply, and the frontier question becomes whether any observable is simultaneously measurable, evasion-resistant and correlated with danger.
Established Meanwhile the instruments remain in the developers’ hands. The only frontier-scale dangerous-capability battery on the public record is Google DeepMind evaluating a Google DeepMind model. Frontier That is not a criticism of the work, which is careful and reports a negative result rather than a favourable one; it is a structural observation about who holds the instruments, with an exact historical analogue in pharmaceutical regulation before the regulator had laboratories.
Frontier Sequestration is the only known repair for contamination, and it converts a benchmark into an institution. If the test set is public it is training data; if it is private it is unauditable by anyone outside the holder; the only escape is a third party who holds the items, administers them and publishes the protocol without publishing the questions. Established That design is old and comes from competitive machine learning and from clinical trials alike: a held-out set with a scoring service, a submission budget per model to stop the leaderboard being reverse-engineered by repeated queries, a documented item-construction method, and a stated refresh schedule. Frontier What that describes is not a benchmark but a testing laboratory with a custody chain, which is the difference between a leaderboard and a certification. Speculative No public evaluator on the roster above operates one at frontier scale on the record this brief could obtain, and the cost of doing so sits an order of magnitude below the declared budgets of the larger institutes, which makes its absence a choice rather than a constraint.
Established The uncertainty problem has a solved statistical answer that the field mostly does not apply. Miller sets out the routine machinery — central-limit standard errors on mean scores, clustered errors where questions arrive in related groups, paired analysis when two models are compared on the same items, and power analysis before the run rather than after it — and the reason that paper had to be written is that published evaluation results commonly report a bare percentage. Frontier A bare percentage cannot support a regulatory decision, because it cannot separate a real difference from a sampling artefact, and the first question a contested conformity assessment will ask is whether the gap between sixty-one and sixty-four per cent survives an interval. Established Reuel and co-authors, scoring benchmarks against an explicit criteria set rather than arguing about them, found the weakest dimensions across the field to be documentation, maintenance and statistical treatment rather than task design. Frontier Gaming then follows automatically. Any score that becomes a threshold acquires an optimisation target, and an instrument with undeclared constructs, decaying calibration and no interval is the easiest possible thing to optimise against without moving the underlying property at all. Speculative The defence is not better benchmarks but audit-style design: sequestration, randomised item selection, adversarial re-testing after certification rather than only before it, and publication of the elicitation protocol so that a score can be reproduced by a second party who did not want it to hold.
Frontier A national sandbox in every Member State, each with different admission rules, is the closest thing to a randomised institutional trial that regulatory science will ever be handed, and nobody has pre-registered an evaluation of it. The states will differ in eligibility criteria, fees, duration, staffing, sectoral focus and exit conditions, and those differences are assigned by domestic administrative politics rather than by the characteristics of the applicants, which is precisely the variation an evaluator wants and almost never gets. Established The design that exploits it is standard and old: agree a common minimum data schema before the first cohort, record every applicant rather than every entrant, and compare outcomes across states whose rules differ. Frontier Recording the rejected applicants is the whole trick, because a sandbox that publishes only its graduates is reporting a treated group with the control group discarded, which is how the ninety per cent figure came to exist. Speculative The window closes at first intake: once the cohorts have been admitted under two dozen incompatible record-keeping regimes the natural experiment is unrecoverable, and the field spends another decade arguing about a number from 2017.
4 · Technological bottlenecks
Frontier The workback target is deliberately harsh: an authority that can independently verify a frontier developer’s capability claims, including the negative ones, with a stated false-negative rate. Verifying a positive claim is easy — if a model can write exploit code, run it and watch. Established Verifying a negative — this model cannot uplift a novice to a working pathogen — is the actual regulatory need, and it is a claim about an unbounded search space. Every link below makes a negative claim more supportable.
Frontier The chain. L1: a legally enforceable pre-deployment access right at white-box level, measured as the count of frontier models where access was granted against models released. Established Not achieved; the record shows the opposite as of April 2024. Frontier L2: a secure evaluation enclave with cleared staff, air-gapped compute and statutory non-disclosure, so L1 is grantable without the developer losing control of its weights, sized in accelerator-equivalents against the developer’s own training cluster. Established No such enclave appears on the record this brief could obtain. Frontier L3: a published elicitation protocol with a measured elicitation gap — best capability from naive prompting against best capability after N hours of expert scaffolding, reported as a curve with a plateau estimate and a confidence interval. This is the first genuinely novel measurement in the chain and nobody publishes it. Speculative L4: a capability ceiling estimator, extrapolating from the L3 curve an upper bound on what any elicitation could extract, with error bars, and calibrated against held-out models. Speculative L5: invariance results — how far that ceiling moves per unit of fine-tuning compute and per tool added. Frontier L6: inter-rater reliability, two institutes evaluating the same model blind, reported as an agreement coefficient with variance decomposed into model, protocol and evaluator. L7: a false-negative rate for the whole battery, established by injecting capabilities into test models and scoring how often the battery finds them. Frontier L8: evaluator capacity at scale, and L9: mutual recognition, so the L8 burden is shared rather than replicated. Speculative L10: an authority holding all of it, with subpoena power and a standing scientific panel, able to contradict a developer’s negative safety claim and make the contradiction stick.
Speculative L4 binds. L1 and L2 are political and financial — hard, but of a kind states routinely solve when motivated, as nuclear safeguards, drug approval and aviation certification all demonstrate. L8 and L9 are scaling problems downstream of a working method. Speculative L4 is different in kind: a scientific result nobody has, that nobody is currently trying to produce as a calibrated quantity, and without which every other link buys activity rather than evidence. Frontier An authority with full weight access, a large enclave and a thousand evaluators, but no bound on what elicitation could extract, can still only ever report we did not find X — a statement whose informational content is exactly zero without L7’s false-negative rate, and whose extension to future elicitations requires L4. Speculative The same missing primitive underlies the hardware-attestation route in L2 and reappears, in a different field, in Cognitive Liberty; that convergence is set out under adjacent technologies below.
Frontier Operational transfer is the sixth property and the one that silently defeats the chain above. Every rung from L3 to L7 measures a model under laboratory conditions, while the regulated object is a deployed configuration with tools, retrieval, memory, a user population and an adversary. Established The nearest thing the field has to an operational measurement is the human-subject uplift trial, in which participants attempt a task with and without model assistance and the output is scored blind by domain experts, and the published examples of that design have reported no statistically significant uplift while stating openly that they were underpowered for the effect sizes that would matter. Frontier Underpowered is the load-bearing word: a trial that cannot detect a doubling of success probability on a rare-event task is not evidence of absence, and reporting it as though it were is the most likely single route by which a measurement failure becomes a policy failure. Speculative Adding a rung below L7 — a powered, pre-registered uplift protocol with a stated minimum detectable effect and a named comparator — would cost less than any of L4 through L7, and it is the only rung that measures the quantity the statute is actually written about.
Frontier The sandbox adds a rung the chain above does not have, and it sits below L1. L0: a stated evidence threshold — what a supervised trial must show before a system graduates from it, written down before the cohort opens and expressed in the same units the eventual conformity assessment will use. Established Regulators in older sectors do this as a matter of course: a medical device pathway specifies the endpoint, the comparator and the analysis plan in advance, and the airlock-style pilot that the United Kingdom medicines and devices regulator opened for artificial intelligence as a medical device inherits that habit from the sector rather than from the technology. Speculative Nothing this brief could find states such a threshold for an AI sandbox, and without one the graduation decision is a judgement call, which means the scaling step from trial to general rule carries no evidentiary content at all: the regulator learns, but what it learns is not written in a form a second regulator can check, contest or reuse.
5 · Research dependencies
Frontier This subject waits on results from four other research programmes, none of which is organised to deliver them. The first is an elicitation science: a body of work that treats the extraction of capability from fixed weights as a measurable process with a cost curve and a plateau. Speculative No such literature exists as a coherent programme; what exists is scattered capability-elicitation folklore inside frontier labs and a growing agent-scaffolding literature that measures new highs rather than bounds.
Frontier The second is metric science. If Schaeffer, Miranda and Koyejo are right that apparent emergence is a property of discontinuous metrics, then the whole practice of capability thresholds needs a theory of which scoring rules are appropriate for which regulatory purposes. That is a statistics problem, it is cheap relative to everything else in this brief, and no regulator has commissioned it. Frontier It connects directly to the instrument problems set out in Intelligence Measurement, which owns the general question of what a capability score means.
Speculative The third is hardware attestation: a chip-level mechanism that can testify, in a tamper-evident and internationally acceptable way, to what computation a device performed. Frontier It is named as a precondition in the compute-governance literature and in the verification literature, and it is a semiconductor engineering programme rather than a policy one. Handwave Nothing in the AI governance literature specifies it to the level a chip designer could build against.
Frontier The fourth is the legal commentary this brief could not read, and it is worth naming as a dependency because it is unusually cheap to close. Established The article-by-article commentary on the Act’s high-risk chapter, the commentary on the EU database for high-risk systems, and the dedicated analysis of the standardisation system all exist as verified bibliographic records. Speculative Reading three documents would settle the staged-application question and the standards timetable that this brief has had to leave open, and would do more for the accuracy of public discussion of the Act than any new empirical work.
6 · Required experiments
Frontier The recommended sequencing is not the logical order of the chain. L4 binds, but L4 is not the first thing to attempt. The real-world order is L3, then L7, then L6, then L4, with L1 and L2 pursued in parallel because they gate everything and have long political lead times.
Frontier Experiment one: the elicitation-gap curve (L3). Take a fixed set of weights and a fixed dangerous-capability task. Measure best performance under naive prompting; then under 1, 4, 16 and 64 expert-hours of scaffolding, tool provision and prompt engineering by a red team blind to the earlier results. Publish performance against expert-hours, with a fitted plateau and a confidence interval. Frontier Nothing in the published literature reports this curve, and it requires no new science — only a protocol, a red team and the discipline to publish the shape rather than the maximum.
Speculative Experiment two, and the highest-value unbuilt instrument in AI governance: a red-team injection benchmark that gives dangerous-capability batteries a published false-negative rate (L7). Construct test models into which specific capabilities have been deliberately introduced, at known strengths and by known routes. Run standard batteries against them blind. Report the fraction of injected capabilities the battery missed. Frontier That single number converts every subsequent negative finding from an uninterpretable statement into evidence with a known error rate. It is cheaper than the capability-ceiling estimator, it is directly fundable by any of the existing institutes, and nothing in the published literature does it.
Frontier Experiment three: blind inter-rater reliability (L6). Two institutes evaluate the same model independently, without conferring, and the findings are compared as an agreement coefficient with variance decomposed into model, protocol and evaluator. Established The network’s July 2025 joint exercise on evaluating AI agents, focused on information leakage and cyber-security risk, is the right shape; no reliability statistic from it is published. Speculative Publishing one would be the cheapest credibility-building act available to the network, and its absence is more informative than most of what is published.
Speculative Experiment four: invariance under modification (L5). Measure how far any bound established in L3 or L4 moves per unit of fine-tuning compute, per tool added, and per scaffolding change. Frontier A certified property of a model is only useful to a product-safety regime if this delta is bounded, and no published work reports it.
Frontier Experiment four, which does not displace experiment one but carries a deadline the others do not: pre-register the sandbox cohort comparison before first intake. Agree a minimum schema across participating states — applicant characteristics, admission decision and stated reason, supervised period, deviations permitted, incidents recorded, graduation decision and post-market outcome — and publish the rejected applicants alongside the admitted ones. Established The analysis is then conventional: a comparison across states whose admission rules differ for administrative reasons, with the applicant pool rather than the entrant list as the frame. Frontier This is the only experiment in this brief whose cost is almost entirely coordination rather than compute, and the only one that expires, because it cannot be run retrospectively: the counterfactual cohort is never recorded unless somebody decides in advance to record it. Speculative Its payoff is the first evidence in the subject about whether supervised trial periods change deployment outcomes, as against whether they change financing outcomes, which is so far the only effect anyone has measured.
7 · Engineering requirements
Frontier The enclave is the concrete build. What L1 and L2 require is unglamorous and specifiable: air-gapped compute sized in accelerator-equivalents against the developer’s own training cluster, cleared staff, physical security, and a statutory non-disclosure regime robust enough that a developer’s counsel will advise handing over weights. The nearest existing analogues are classified national-laboratory computing facilities and clinical-trial data safe havens, both of which took a decade to become routine. Established Nothing resembling it appears in the institute roster, the budgets or the published methodology on the record this brief could obtain.
Speculative Hardware attestation is the second build, and it is a semiconductor programme. The requirement is a mechanism at the accelerator level that produces a tamper-evident attestation of what computation was performed, acceptable to a foreign inspectorate. Frontier The compute-governance literature names concentration in the supply chain as the reason this is feasible at all; the verification literature groups its most powerful techniques under “requires advanced hardware controls”. Handwave Neither specifies the threat model against a determined nation-state adversary with physical possession of the device, which is the case that matters and the hardest one in trusted computing.
Frontier Third: agent logging that is tamper-evident against the operator. Agent identifiers, real-time monitoring and activity logging are the proposed primitives. Speculative The engineering difficulty is not the logging; it is that the party with the strongest incentive to alter a log is the party that runs the infrastructure, which puts the requirement in the same class as financial audit trails and requires the same kind of cryptographic and institutional scaffolding.
Speculative Fourth, and the cheapest: data-centre signature monitoring. Unauthorised data centres are one of the two violation classes the verification literature targets, and large training facilities have thermal, electrical and satellite-visible signatures. Frontier The engineering is largely solved by existing remote-sensing and grid-monitoring practice; what is missing is a declared-facility baseline against which anomalies could be scored. Handwave Without the baseline, the false-positive rate is unbounded and the technique cannot survive its first diplomatic use.
8 · Adjacent technologies
Speculative The strongest cross-brief finding in this cluster: AI governance and cognitive liberty converge on the same missing primitive, and neither literature appears to know it. Both reduce, at the binding step, to verifiable claims about what computation was performed. Frontier An authority trying to check a frontier developer’s capability claim needs to know what ran on which hardware, on what data, for how long — that is the enclave requirement and the hardware-attestation requirement stated in a single sentence. A person trying to establish that their neural data was decoded without consent faces the mirror image: the decode leaves no trace on the subject’s side, so the only possible evidence is a record of the computation the other party performed. Speculative Neither field can solve its binding problem without that primitive; the AI-governance literature treats it as a compute-governance implementation detail and the neurorights literature does not raise it at all. See Cognitive Liberty, where the same gap appears from the other side.
Frontier A second shared shape. Three separate subjects in this cluster need a bound over a search space and have only samples from it: the capability ceiling under unbounded elicitation here; whether a decoder could be made to work without cooperation in cognitive liberty; and whether a collective could beat its best member under an optimal decomposition in Collective Intelligence. Speculative That is a methodological problem, not three separate empirical ones, and it is unowned.
Frontier The other live adjacencies: Multi-Agent Intelligence Systems, which owns the individuation question that agent identifiers presuppose; Artificial General Intelligence, whose capability-forecasting disputes set the urgency parameters of every instrument here; Future Legal Systems, which owns the liability route this brief argues cannot carry the tail; and Ultra-Efficient Computing Energy Systems, because a compute threshold denominated in floating-point operations is a moving target if the energy and hardware cost per operation keeps falling.
9 · Institutional requirements
Established The capacity question, answered numerically. Two staff counts are on the public record: roughly 23 in Japan and at least 30 in South Korea. Frontier A professional evaluator labour market of the size the chain above implies would need low thousands of qualified people, and there is no training pipeline that produces them. Speculative The nearest analogues — clinical trial statisticians, nuclear safeguards inspectors, aviation certification engineers — each took two to four decades and a dedicated professional-formation route to build, and each has a defined body of knowledge that AI evaluation does not yet have.
Established Institutional durability is measurable and the measurement is bad. The UK institute was renamed to AI Security Institute in February 2025 and the US institute became the Center for AI Standards and Innovation in June 2025, with the mission reframed toward innovation over safety considerations; observers read the UK change as signalling a shift away from ethical concerns. Frontier Two of the three best-resourced public evaluators were re-scoped inside sixteen months. Speculative Whatever else that is, it is data about durability, and durability is a precondition for accumulating evaluation expertise: a body that is re-mandated every eighteen months cannot run a ten-year measurement programme.
Established The coordination layer exists in outline. The network of institutes was formed at the AI Seoul Summit in May 2024 with the UK, US, Japan, France, Germany, Italy, Singapore, South Korea, Australia, Canada, the European Union and Kenya as members; a UK–US agreement for joint safety testing dates from April 2024; a joint exercise on evaluating AI agents ran in July 2025; and members met at NeurIPS 2025. Frontier Mutual recognition of evaluations — the mechanism that would let one jurisdiction’s work count in another — has begun in form and not in substance.
Established Inside the EU, the enforcement design is a two-part split. The AI Office is attached to the European Commission, coordinates implementation across Member States and oversees general-purpose provider compliance; a scientific panel of independent experts supplies technical advice and supports enforcement of the general-purpose rules. Frontier Set that against the International AI Safety Report and the split becomes clear: the Report is the assessment half without the enforcement half, and the AI Office with its panel is the enforcement half without a demonstrated measurement capacity. Speculative Nobody currently holds both, and the whole regime is a wager that the halves can be joined faster than capability advances. Frontier The comparative institutional question — whether an advisory panel can ever discipline an executive that funds it — belongs to Scientific Advisory Institutions and is not re-argued here.
Established The sandbox is the one governance instrument in this subject that was designed as an experiment, and it is about to be run at scale. The European regime requires national AI regulatory sandboxes: controlled environments in which providers develop, train, test and validate systems under direct regulatory supervision before the system is placed on the market, with participation conditions and a supervisory contact settled in advance. Frontier This brief takes that description from the Commission service desk’s own guidance and prints no article number for it, consistent with its treatment of the rest of the instrument. Frontier The obligation is widely reported to fall due in August 2026, which would place a sandbox in every Member State within a year of this brief; the date is flagged rather than asserted, because it depends on the same Article 113 allocation this brief declined to compute earlier. Established The instrument itself is not new. The United Kingdom financial conduct regulator opened the first one in 2016, and its own lessons-learned report a year later described roughly ninety per cent of first-cohort firms as proceeding toward wider market launch.
Frontier That ninety per cent is the exact shape of the evidence problem. It is a self-reported figure about a small, discretionarily admitted cohort with no control group, and it has been quoted for a decade as though it demonstrated that sandboxes work. Established The one quasi-experimental estimate this brief knows of is Cornelli and co-authors at the Bank for International Settlements, who compared sandbox entrants against matched non-entrants and found entry associated with a materially higher probability of raising capital and a substantial increase in the amounts raised. Frontier Read carefully, that result is about financing rather than about safety: the measured effect of the sandbox is a certification signal to investors, which is a real effect and not the one the instrument is justified by. Speculative If the dominant effect of an AI sandbox is also a signalling effect, then admission becomes a scarce and valuable licence issued at regulator discretion, and the capture risk changes shape: the danger is not that firms lobby for weak rules but that they compete for a place in the queue, and the regulator acquires an interest in the success of the firms it has admitted. Established Liability is not suspended by any sandbox this brief is aware of — participants remain answerable to third parties, and the European regime pairs the sandbox with real-world testing provisions that turn on informed consent rather than on immunity. Frontier The consequence is that a sandbox cannot observe the one thing a regulator most needs to see, which is how a system fails when nobody is watching it.
10 · Ethical & societal considerations
Established The strongest objection to compute governance comes from its own proponents. The authors who set out compute as a governance lever warn in the same paper of privacy violations, economic harm and power centralisation. Speculative Citing the concession rather than the critic is both more persuasive and better sourced, and it makes the shape of the trade clear: an observable that is detectable and excludable enough to regulate is by construction an observable that can be used to decide who is permitted to compute.
Frontier Agent logging has a symmetrical problem. The visibility measures proposed for AI agents — identifiers, real-time monitoring, activity logging — are described by the people proposing them as carrying privacy and power-concentration costs. Speculative A comprehensive agent log is a record of what people asked machines to do on their behalf, which is close to a record of what they intended, and the governance value of the log rises with exactly the property that makes it dangerous.
Frontier Publishing an evaluation protocol can publish the capability. The requirement that evaluations be reproducible collides with the requirement that dangerous-capability methods not be a manual. Speculative No published resolution exists; the nearest workable proposal is a trusted-third-party enclave with cleared staff and legal immunity from disclosure, which shifts the problem into institutional design rather than solving it.
Frontier The burden of proof is currently on the wrong party and everyone knows it. A regime in which the developer runs the evaluation, sets the protocol and reports the result is one in which a negative finding is a claim by an interested party. Speculative The honest response is not distrust but adversarial reproducibility: protocols published, replication possible without weights or inside an enclave, and inter-rater statistics reported. Established None of the three is in place.
Speculative And the correct governance posture may be heterogeneous, which nobody wants. If adaptation dominates control for some harms and not others, then the right regime differs per harm — hardening for deepfake provenance, interdiction of physical inputs for biological risk, something else again for cyber. Frontier That conclusion is unpopular on both sides of the debate: it denies the safety camp a unified licensing regime and denies the acceleration camp a clean exemption.
11 · Civilizational implications
Speculative The sharpest claim in this subject is a conditional, and the current regime is a bet that its antecedent is false. The conditional: if there is no observable that is simultaneously measurable, evasion-resistant and correlated with danger, then frontier AI is ungovernable by measurement, and the only remaining route is liability. Frontier The liability route then requires three things: causation traceable from a harm back to a model, which needs provenance infrastructure that does not exist; harms that are compensable, which fails for catastrophic and irreversible harms by construction; and developers solvent against the tail, which no insurer currently prices. Speculative The second binds absolutely. Liability is a post-hoc instrument and the tail risk is the entire motivating case. Handwave So if the antecedent holds, neither route works — and the honest description of every regime now in force is that it is a wager on the antecedent being false, placed without anyone having tried to test it.
Speculative There is a genuinely surprising counterweight, and it reverses the usual assumption. The two violation classes that international verification targets — very large training runs and unauthorised data centres — are physically large, with thermal, electrical and satellite-visible signatures, which is exactly why arms control worked. Frontier What blocks verification is not physics but a declared-facility baseline, which requires agreement to declare. Speculative That means the verification problem may be more tractable than the treaty problem, which is the reverse of the assumption that organises most commentary. Handwave If true, the highest-leverage civilizational move in this subject is diplomatic rather than technical, and it looks like the 1963 and 1968 precedents rather than like a research programme. The comparative case for and against that reading sits in Global Cooperation Models.
12 · Timelines
These horizons track measurements, not products. Each entry names a specific instrument or statistic that either exists by that date or does not.
- 10 yr: Frontier The harmonised standards under the EU regime either exist and are cited in conformity assessments, or the high-risk tier remains formally in force and operationally empty; that single fact will settle more about the Act’s effectiveness than any amount of doctrinal argument. Frontier At least one dangerous-capability battery is published with a measured false-negative rate, or none is, and if none is then the entire decade of negative safety findings remains uninterpretable. Speculative A blind two-institute reliability exercise reports an agreement coefficient, and pre-deployment white-box access becomes a legal entitlement in at least one jurisdiction or remains a voluntary courtesy everywhere.
- 25 yr: Speculative Either an elicitation-gap curve with a fitted plateau is a routine deliverable in frontier model documentation, or the field has concluded that elicitation is unbounded in practice and the product-safety framing is abandoned for a control-and-adaptation framing. Speculative Hardware attestation is either deployed at accelerator level and accepted across at least two blocs, or the compute lever has been retired as unenforceable. Frontier The governable unit has settled: either a model registration regime or an agent-identifier regime dominates, and the loser’s instruments become vestigial. Speculative Evaluator headcount is either in the low thousands with a professional formation route, or it is still in the low hundreds and the capacity question is answered in the negative.
- 50 yr: Speculative A capability ceiling estimator with a published calibration record either exists or has been shown to be impossible, and either result reorganises the subject. Speculative Mutual recognition covers a majority of frontier releases, or evaluation has fragmented along bloc lines and models are certified for jurisdictions rather than for use. Handwave A standing international body with inspection rights and an enclave exists, on the safeguards model, or the absence of one is the accepted state of the world and the governance literature has moved wholly to adaptation.
- 100 / 250+ yr: Handwave The question at this horizon is not which instrument won but whether measurement-based governance was ever the right frame. Handwave The optimistic branch has a mature evaluation science with error bars, a professional inspectorate and a treaty regime, and looks in retrospect like nuclear safeguards: slow, imperfect, sufficient. The pessimistic branch is that no bound over elicitation was ever found, liability failed on the tail exactly as predicted, and what governed was whatever hardening the defended surfaces happened to receive. Handwave Cheerfully far from evidence, and stated because the alternative — assuming the instruments will arrive because they are needed — is the assumption every governance document in this brief quietly makes.
Frontier One measurement-science milestone belongs on the ten-year line and is not in the list above: either at least one sequestered, third-party-administered dangerous-capability test set is operating with a published refresh protocol and a per-model submission budget, or every capability finding of the decade rests on instruments whose calibration decayed on publication. Speculative The same date settles whether reported evaluation results carry intervals as a matter of course; the statistics involved are undergraduate, so the failure mode there is not difficulty but the absence of anyone obliged to apply them.
13 · Technology tree & dependencies
- Depends on This brief depends on the capability-forecasting work carried in Artificial General Intelligence, because every instrument described here is calibrated to an implicit forecast. A compute threshold set at a fixed value is a bet on the capability-per-operation curve; a staged application timetable is a bet on how far capability moves during the staging; a voluntary code with a mandatory shell behind it is a bet on how long the shell has to hold. It also depends on Intelligence Measurement for the underlying question of what a capability score is a score of, since the metric critique that undercuts emergence as a governance premise is a special case of the general instrument problem that brief owns. Neither dependency is thematic. If the capability-per-operation curve steepens, the 1025 trigger stops discriminating within a few model generations and the systemic-risk tier either empties or fills indiscriminately, and no amount of institutional design compensates.
- Requires (not on this map) Both constraints are institutional and both are measurable, which is why they are stated as constraints rather than as concerns. The first is not a staffing shortfall but a structural one: the expertise required to evaluate a frontier system sits disproportionately inside the firms being regulated, the only frontier-scale dangerous-capability battery on the public record was run by a developer on its own model, and the two public staff counts available are roughly 23 and at least 30 people against a requirement in the low thousands. The second is a dated one: the high-risk tier acquires operational content only when CEN/CENELEC Joint Technical Committee 21 delivers harmonised standards and those standards are published, after which compliant products are presumed to conform. This brief could not obtain a delivery timetable, a deadline or any record of delay, and it declines to assert one; the gap between an obligation’s date and the standard that operationalises it is named here as an open question with a specific closing source rather than filled with an impression. A third constraint is added by this pass and is of the same kind, administrative rather than scientific: the national sandboxes required under the European regime will generate the only cross-jurisdictional variation in AI supervision that anyone is going to get, and that variation is legible only if each sandbox records and publishes its applicant pool rather than its graduates. Keeping that record is not research and needs no new capability; it is a decision that has to be taken before the first cohort is admitted, after which the comparison is unrecoverable and the instrument reverts to a collection of testimonials.
- Enables What this brief would enable, if its binding link were delivered, is unusual: a calibrated capability ceiling with a stated miss rate would be the first piece of evidence in the subject that could support a negative safety claim, and negative safety claims are what every downstream instrument silently assumes it has. Insurance pricing for frontier deployment, mutual recognition between jurisdictions, safety cases built on capability absence, and any treaty with an inspection regime all currently rest on a quantity nobody has measured. Nothing on this map waits on that quantity today, and that absence is itself the finding rather than an oversight: the measurement sits between machine learning, metrology and regulatory science, and is owned by none of them, which is exactly the profile of the cheap experiment that nobody runs.
- Adjacent Adjacent work sits in Cognitive Liberty, which needs the same missing primitive from the opposite direction and does not know it; in Multi-Agent Intelligence Systems, which owns the individuation problem that agent identifiers presuppose and cannot resolve on their own; in Future Legal Systems, which owns the liability route this brief argues cannot carry catastrophic and irreversible harms; and in Global Cooperation Models, which owns the declared-facility baseline problem that turns out to gate international verification more tightly than any technical obstacle does. The general theory of regulatory institutions is carried across Category VIII and is deliberately not restated here; what this brief adds to it is the measurement layer that the political theory assumes and does not supply.
14 · Common misconceptions & speculative claims
Established “The EU AI Act came into force in 2024 and applies from 2026, and it was passed on 21 May.” Entry into force on 1 August 2024 is right; a single 2026 application date is not, because obligations stage over 6 to 36 months and the general-purpose model provisions bit on 2 August 2025. Established And 21 May 2024 is Council adoption, not the citation: the instrument itself bears 13 June 2024, the Official Journal published it on 12 July 2024, and Parliament voted in March. Frontier Five defensible answers to “when was it passed”, of which press coverage reliably picks the second-least legally significant. Frontier The staging matters for the same reason: prohibitions, general-purpose regime and high-risk regime are three regulatory events with three readiness states, and one date hides all of it.
Established ENTHUSIAST-SIDE: “The AI Act regulates AI.” It regulates uses, tiered by risk, plus a separate model-level regime for general-purpose systems, and most AI falls in the minimal-risk tier with no obligations at all. Frontier Campaigners who call it the world’s first comprehensive AI law should be pressed on the word comprehensive, which the four-tier structure specifically disclaims.
Established ENTHUSIAST-SIDE: “The Act bans dangerous models above 1025 floating-point operations.” It does not. Crossing the threshold moves a model into the systemic-risk category of the general-purpose regime, which brings obligations, not prohibition. Established The prohibitions attach to practices — the sources reachable here name emotion recognition and real-time remote biometric identification, with law-enforcement exemptions — not to model scale. Speculative The conflation inflates public expectations of what the Act does to frontier labs.
Established “There are AI Safety Institutes in the US and UK.” The UK body has been the AI Security Institute since February 2025 and the US body the Center for AI Standards and Innovation since June 2025, the latter reframed toward innovation over safety considerations. Frontier Using the old names in the present tense names bodies that no longer exist under them, and the renaming is itself evidence on the durability question.
Frontier “The institutes test frontier models before release.” On the record this brief could obtain, no pre-deployment evaluation reports are recorded, and as of April 2024 many companies had not shared pre-deployment access to their most advanced models. Established What is published is a framework — Inspect, May 2024 — and a joint agent exercise in July 2025. Speculative The distance between what these bodies are assumed to do and what they have published is the largest expectation gap in the subject, and it is not the institutes who created it.
Frontier “A dangerous-capability evaluation that finds nothing is reassuring.” It is uninterpretable without a false-negative rate, and no source reachable here publishes one for any battery. Established The only worked frontier-scale battery frames its own negative result as groundwork in preparation for future models rather than as a safety conclusion. Frontier This is the most important misconception in the topic and it is held by sophisticated people, including people who commission evaluations, and the remedy is one statistic and a protocol.
Frontier “Emergent capabilities mean regulators must catch sudden jumps.” Contested at the root: Schaeffer, Miranda and Koyejo argue emergence appears due to the researcher’s choice of metric rather than to changes in model behaviour with scale. Speculative If they are right, the jump-catching rationale for compute thresholds weakens considerably and the correct posture shifts from tripwires toward continuous monitoring of a smooth curve. The debate is live; this brief does not adjudicate it, but notes that the governance literature has largely not absorbed it.
Established “Compute governance is a backdoor to controlling who can compute.” The strongest version of the worry is not the critics’ but the proposal’s own authors’, who warn in the same paper of privacy violations, economic harm and power centralisation. Speculative Quote the concession rather than the critic; the honest position is that the worry is correct and the trade may still be worth making, which is harder to say than either slogan it replaces.
Frontier “International agreement on AI is impossible to verify.” Ten techniques are enumerated in the published literature, each with historical analogues and circumvention analysis, against two physically large violation classes. Speculative Verification is partially specified and untested, not impossible, and the more surprising position argued above is that it may be easier than the treaty.
Speculative “Voluntary commitments are worthless.” Too strong. The general-purpose AI Code of Practice of 10 July 2025 is voluntary, but sits inside a regime carrying penalties up to 7 per cent of worldwide turnover for the mandatory parts, and a voluntary instrument inside a mandatory shell behaves differently from one standing alone. Established The correct critique is about measurement, not voluntariness: nobody can tell whether the commitments are being met, and this brief could not even establish who signed.
Handwave “No existing institutional form can govern frontier AI.” The strongest version deserves stating because it is the exotic end of the subject and it is not silly. Product-safety law presumes a fixed artefact with fixed properties, and a model that is fine-tuned, scaffolded and given tools after certification is not that. Process regulation presumes the supervisor can read the books; the audit literature says the supervisor cannot read the weights. Nuclear safeguards work because fissile material is conserved and countable, and weights are copyable at zero marginal cost. Speculative For governance to succeed anyway, three things must hold: a certified property must exist that is invariant under fine-tuning and scaffolding; it must be cheap to re-verify after each modification; and verification must be possible without handing over the weights. Frontier The first binds and none is demonstrated. Speculative The mainstream reply is that the same was said of pharmaceutical regulation before a regulator with laboratories existed, and that such institutions are built by first requiring the evidence and then discovering how to produce it. Handwave That is a historical analogy rather than an argument, but it is the best one available and it has been right before.
Speculative A minority position worth taking seriously: the measurement problem is overstated because dangerous capabilities are conjunctive. Catastrophe requires capability and opportunity and tacit knowledge and materials and the absence of interdiction, and conjunctions are easier to break than to build. Frontier A regulator who cannot measure capability might still reliably measure opportunity. Speculative For that to work the narrowest link must be physical rather than informational and independently regulable — arguably true for biological risk, which is why nucleic-acid synthesis screening is plausibly the field’s highest-leverage single intervention, and arguably false for cyber. Frontier See Synthetic Biology for the chokepoint argument in its own setting.
Handwave And the furthest-out coherent claim: the first real AI regulator will be an AI. The volume of models, deployments and agents will exceed human review capacity by orders of magnitude, and the control literature already demonstrates a weaker trusted model monitoring a stronger untrusted one as a working protocol. Speculative For that to become governance rather than an experiment, three things must hold: the trusted-monitor asymmetry must survive at scale, having been tested only in a narrow programming domain; the monitor’s findings must be legible to humans, or the regulator has merely relocated the black box; and legal systems must accept machine-generated findings as evidence. Frontier The first is an open empirical question of enormous importance: if monitoring is easier than generating, governance scales with capability; if not, it does not scale at all, and everything in this brief describes the last period in which the question was still open. Handwave Nobody has run the experiment that would say which.