1 · Concept overview

Established AI-biology governance is the machinery that sits between a model that knows biology and a bench that can make something. It has five layers, and they are institutionally separate: evaluation of what a frontier model can do in the biological domain; safeguards built into the deployed system; governance of who is granted ungated access; screening of synthetic nucleic acid orders at the point of manufacture; and auditability, meaning whether anyone outside the organisation being governed can check any of the four layers above. This brief owns those five and nothing else.

Established The brief withholds capability detail by design, and the omission is load-bearing rather than decorative. Nothing below names a sequence, an agent, a specific screening blind spot, or any step of a route from a model output to a physical product. Where the underlying literature contains such material — and some of it does — the source is cited and only the governance result is reported. A reader looking for a gap map will not find one here. The scientific question of what generative models can design, and whether those designs work in a wet lab, belongs to a sibling commission on generative biomolecular design; the fabrication and editing capabilities themselves are owned by Synthetic Biology and Genetic Engineering, and are referenced, not repeated.

Frontier The stack is measured upside down: the layer closest to the bench has the best numbers, and the layer that triggers every other control has the worst. Nucleic acid synthesis screening has a public benchmark dataset, a monthly proficiency programme running since August 2025, published sensitivity medians, and a documented live-order stress test. Frontier-model biological capability evaluation — the thing that causes a developer to switch safeguards on at all — has benchmark scores, a contested relationship between those scores and real-world uplift, and no published false-negative rate for any deployed biological safeguard stack this brief could obtain. The best-instrumented control in the chain is being asked to compensate for the worst-instrumented one, and nobody has measured whether it can.

Established Four words in this subject mean four different things and are routinely conflated. Screening is what a synthesis provider does to an order: match it against sequences of concern, and verify the customer. Safeguards are what a model developer builds into a deployed system: refusal training, classifiers, account-level enforcement. Access governance is the decision about who is allowed past those safeguards on purpose, and on what evidence. Auditability is whether a party with no commercial interest can verify any of the first three. Each has a different institutional home and a different failure mode, and a policy argument that slides between them is usually wrong.

Established The boundary against the neighbouring briefs is drawn at the instrument, not the theme. AI Governance owns the general claim that AI statutes delegate the hard question to an evaluation science that does not exist; this brief does not re-argue it. Existential Risk Governance owns treaty-level machinery, including the Biological Weapons Convention and the finding that no body in that family has a measured counterfactual. What is added here is the one place where those two literatures collide with a working technical control: an order-screening regime that predates modern AI by fifteen years, has real instrumentation, and is now being asked to hold a line it was not designed for.

2 · Current scientific position

Established There is no United States mandate to screen synthetic nucleic acid orders, and as of September 2026 there has not been one for sixteen months. The 2023 Department of Health and Human Services Screening Framework Guidance remains the operative document and is advisory. The 2024 Framework for Nucleic Acid Synthesis Screening, published by the Office of Science and Technology Policy on 29 April 2024, worked by procurement condition rather than by regulation: recipients of federal life-science funding were to buy only from providers attesting to compliance, with the obligation attaching from 26 April 2025. Executive Order 14292 of 5 May 2025, Improving the Safety and Security of Biological Research, directed agencies to revise or replace that Framework within 90 days. The 90 days expired in August 2025. The replacement policy issued on 28 July 2026 addressed high-risk life-sciences research and gain-of-function funding; it did not revise the synthesis screening framework, which the responsible agency still lists as under revision.

Established The screening tools themselves are the best-measured artefact in AI-biology governance, and they measure well. A National Institute of Standards and Technology test dataset of 999 anonymised 200-base-pair fragments — 249 bacterial and 250 viral true positives against 250 bacterial and 250 viral true negatives — was screened by six independently developed tools (Aclid, the Common Mechanism, FAST-NA Scanner, SeqScreen, SecureDNA and UltraSEQ). Individual tools reached sensitivity at or above 95 per cent and accuracy at or above 97 per cent; 92.3 per cent of sequences drew a unanimous verdict across all six; a majority vote across the six classified every sequence correctly. From August 2025 NIST has run this as a standing monthly programme, sending providers 1,000-sequence sets composed of 200 true positives, 200 true negatives and 600 ungraded sequences. As of July 2026 the reported provider median was a sensitivity of 0.9675 and an accuracy of 0.9788, against pass thresholds of sensitivity above 0.95 and accuracy above 0.75.

Frontier Those numbers measure tools on a curated benchmark, not providers on live commerce, and the one published live test came out differently. In June 2025 NIST ordered three plasmids designed to carry polymerase chain reaction target regions for detecting and discriminating mpox clades, variola virus and orthopoxviruses generally — detection reagents, not agents. Of the twelve orders placed, three were processed without any follow-up contact, and the exercise could not distinguish between three explanations for that: the provider did not screen; the provider screened and judged the sequences safe; or the provider flagged a sequence of concern and cleared NIST as a legitimate customer. That ambiguity is the finding. A benchmark score is a property of an algorithm; whether an order is stopped is a property of an organisation, and only one of the two is instrumented.

Established Coverage is the other unmeasured dimension, and every published estimate of it is a self-report. The International Gene Synthesis Consortium, the main industry body operating a harmonised screening protocol, is estimated to cover roughly 80 per cent of global commercial synthesis capacity, leaving about a fifth outside any voluntary regime. Congressional testimony in December 2025 put the number of providers committed to voluntary screening at more than 30, mostly United States companies. Benchtop synthesisers, which move synthesis inside the customer’s own building, are covered by guidance asking manufacturers to build screening in before synthesis occurs, and by export controls, but by no domestic pre-market security certification this brief could identify.

Established The AI-specific challenge to screening has been demonstrated, disclosed and patched, in that order, and the exercise is the best-run piece of governance in the field. Wittmann and colleagues, publishing in Science on 2 October 2025 with authors from Microsoft, IBBIS, Twist Bioscience and academic groups, red-teamed nucleic acid screening software against variants produced by open-source generative protein design tools and found that redesigned sequences were not reliably detected by the tools as they then stood. The finding was handled as a coordinated vulnerability disclosure: patches were developed with the screening tool maintainers and deployed, and detection was reported as greatly improved. What was not published, and what the governance question needs, is a post-patch false-negative rate against an independently generated adversarial set. The patch is documented; its residual is not.

Established Two free screening systems now carry most of the world’s non-commercial screening capacity, and their published engineering is considerably more serious than the policy around them. SecureDNA, described by 31 authors in National Science Review on 16 February 2026, screens by exact match to short subsequences unique to controlled genes rather than by similarity, which removes the expert-review step that similarity screening requires. Reported figures: screening in under one second for a typical order, about 1,100 nucleotides per second per CPU thread, capacity above four trillion nucleotides a year, zero false alarms from known sequences and 30 false positives arising from unsequenced viral strains — roughly one per 5,000 orders — across 67 million nucleotides of real synthesis from providers in the United States, Europe and China, of which 0.272 per cent would have been recommended for denial. Its governance design is worth naming: five geographically separated keyservers under a three-of-five threshold, biweekly key rotation, a timestamped signature certifying which database version an order was screened against, and privacy-preserving detection logs that allow a split order across several providers to be noticed. Recurring cost to a provider is quoted at roughly $5,000 to $8,000 a year in hardware plus $15,000 to $30,000 a year in colocation (programme’s own figures).

Frontier The other free tool is one funding cycle from maintenance mode, which is the single most concrete fragility in the entire control stack. The Common Mechanism, maintained by the International Biosecurity and Biosafety Initiative for Science, is open source in both code and databases, screens a 200-base-pair sequence in under a second, reports under 2 per cent false positives on engineered sequences, and reached version 2.0 in July 2026. It is explicit about what it cannot do: sequences below 50 base pairs, protein sequences directly, oligonucleotide orders, and sequences with ambiguity codons. It has no third-party audit or certification. Its maintainers state that the tool has had no dedicated funding since the organisation was founded in 2023, ask for $206,000 to $420,500 for one year of development, and say that without it the project shifts into maintenance mode in November 2026. A global screening floor resting on a tool whose annual budget is smaller than a single frontier evaluation contract is a fact about the world, not a rhetorical flourish.

Established On the AI side, the measured record is a set of benchmark scores whose relationship to real-world capability is exactly what is in dispute. The Virology Capabilities Test, built by SecureBio with the Center for AI Safety and published in April 2025, comprises 322 multimodal questions written by 68 contributing expert virologists out of 184 registered, and is scored against virologists answering only within their own sub-speciality. On that instrument the leading model scored 43.8 per cent against an expert in-speciality baseline of 22.1 per cent, reported as outperforming 94 per cent of expert virologists. Frontier Against that, the RAND red-team study of January 2024 found no statistically significant difference in the viability of large-scale attack plans produced with or without model assistance, and its authors said the design lacked the sensitivity to detect the difference they sought. The June 2026 open letter calling for mandatory screening, signed by more than 100 people including the chief executives of the largest model developers and a Nobel laureate in chemistry, concedes in its own text that “the evidence about what this means for present-day biosecurity threats is genuinely mixed”.

Established Both leading developers have crossed their own biological thresholds, and both crossed them on inability to rule the risk out rather than on positive evidence of capability. Anthropic activated its AI Safety Level 3 standard for Claude Opus 4 in May 2025 on the stated basis that it could not rule out the need for it, citing superior performance on proxy tasks and qualitatively different red-team reports, while noting that dangerous-capability evaluation is inherently hard and that it remained uncertain about the length of model access required to produce uplift, the number of potential threat actors, and the complexity of the threat pathway. OpenAI treats its April 2026 flagship as High capability in the biological and chemical domain under its Preparedness Framework. Frontier The safeguard descriptions in both cases are categorical rather than quantitative: constitutional or safety-reasoning classifiers, account-level enforcement, verified-access exemptions, egress bandwidth limits, two-party authorisation for weight access, binary allowlisting, and a bug bounty that turned up a small number of potentially effective jailbreaks. Neither published document gives a false-negative rate, a classifier recall figure, or a bypass rate for the biological stack, and in the OpenAI case the cyber section of the same document is quantified where the biological section is not.

Frontier Third-party evaluation of those safeguards exists, is voluntary, and publishes almost nothing in the biological domain. External groups did get real access in the most recent cycle: SecureBio evaluated pre-release checkpoints between 2 and 9 April 2026 with API-level and system-level biological content filters disabled, and the United States Center for AI Standards and Innovation was given both a launch checkpoint and a reduced-refusal checkpoint and reported that its testing did not indicate a broad increase in national-security-relevant biological capabilities. Established That agency’s own description of its work is built on voluntary agreements with developers, and its published evaluation output to date concerns competitor models and cyber and agent security; no biological evaluation report from a national institute appears on the public record this brief could obtain. The European instrument is similarly shaped: the General-Purpose AI Code of Practice, published 10 July 2025 with obligations biting from 2 August 2025, names chemical, biological, radiological and nuclear risk as one of four mandatory systemic risks, requires independent expert evaluation where appropriate and model reports to the AI Office, but requires publication only of summaries and only in limited conditions.

Established The audit-access literature has been unambiguous since 2024 about why this level of access cannot settle the question. Casper and colleagues, at the 2024 ACM conference on fairness, accountability and transparency, set out the access ladder — black-box query access, grey-box access to activations or sampling probabilities, white-box access to weights and fine-tuning, and outside-the-box access to methodology, training data and internal findings — and showed that black-box testing cannot establish the absence of a dormant capability, because gradient-guided attacks and fine-tuning routinely surface behaviour that unguided querying does not. They pair that with the institutional precedent that the analogous inspection regime in nuclear safeguards employs several hundred inspectors with a legal right of entry. Frontier Two partial technical answers exist: privacy-preserving evaluation using secure enclaves, which in one documented case produced a cryptographic certificate that only evaluation code approved by both the developer and the United Kingdom safety institute had been run, and a cooperative research agreement signed by CAISI in March 2026 to develop that tooling further. Handwave The step that is still assertion is that a developer which retains approval rights over which questions may be asked can be audited at all in the sense the word carries in financial or nuclear practice.

Frontier Only one part of this subject has been subjected to a published quantitative cost-benefit analysis, and it is the oldest part. A December 2025 analysis for the United Kingdom, following Treasury appraisal guidance, estimated that mandatory screening of synthetic nucleic acids above 50 base pairs returns about £3.50 for every £1 spent, with net benefits near £150 million a year domestically rising towards £970 million under international harmonisation, and recommended legislation be proposed by the fourth quarter of 2026. Handwave The benefit side rests on an assumed reduction in the probability of an event that has not happened — the same unobservable denominator Existential Risk Governance identifies across every prevention body it examines. The analysis is careful and its central assumption is unfalsifiable; both are true at once.

3 · Frontier questions

Frontier Does a benchmark score predict operational uplift? This is the question the entire regulatory structure is balanced on and it has no accepted answer. A high score on a virology troubleshooting instrument establishes that a model holds and can apply a body of technical knowledge; it does not establish that a person with that knowledge and no tacit skill, materials, or organisation gets closer to an outcome. The 2026 International AI Safety Report, chaired by Yoshua Bengio with more than 100 experts nominated by over 30 countries and international organisations, names this the evaluation gap and states that it remains difficult to assess the degree to which material barriers still constrain actors. Nothing in the last two years has closed it.

Frontier What is the correct unit of regulation: the model, the task, or the order? Compute thresholds regulate models, capability thresholds regulate behaviours, and screening regulates physical orders. The three do not nest. A December 2025 congressional witness argued that there is no comprehensive shared understanding of what kinds of biological capabilities pose consequential risks, and asked Congress to give a standards agency the authority to define the narrow subset of biological AI models that warrant oversight. Speculative The strongest version of the order-level argument is that the physical chokepoint is the only layer where compliance is observable, and therefore the only layer where verification is possible even in principle.

Frontier Is exact-match screening structurally more robust than similarity screening as design tools improve, or merely differently brittle? The exact-match architecture removes human review and false alarms from known material; the similarity architecture generalises to variants the exact-match index has never seen. The October 2025 red-team result cuts against similarity screening as it stood and was answered with patches; no published experiment has tested both architectures against a common adversarial set generated by a party with no stake in either. That experiment is cheap and nobody has run it.

Frontier How long does a chokepoint stay a chokepoint? Screening governs centralised commercial manufacture. Benchtop devices, cloud laboratories and the roughly one-fifth of capacity outside the voluntary consortium each erode the assumption that orders pass a small number of observable gates. A 2026 perspective in EMBO Reports frames this as an embargo window: a period after a capability is recognised and before it diffuses, in which defenders can act. Whether that window is years or months is unmeasured, and those authors offer no quantitative diffusion estimate.

4 · Technological bottlenecks

Established There is no ground truth for a false-negative rate in either half of the stack, and the two halves lack it for different reasons. For synthesis screening, a full false-negative measurement requires a validated set of things that should be caught, which is precisely the material that cannot be published or widely circulated; the NIST benchmark is a curated proxy, and its maintainers treat it as one. For model safeguards, the equivalent measurement requires an adversary budget and a definition of success, and no deployed biological safeguard stack has published either. This is the deepest technical bottleneck in the subject and it is not obviously solvable by spending more money.

Frontier Any public benchmark degrades as an instrument the moment it is public. The virology instrument is public, which is what makes it a common reference in model cards across at least five developers; the same publicity makes contamination and targeted optimisation possible. The response has been private companions — non-public capability benchmarks licensed to frontier developers rather than released — which restores the instrument and destroys the public’s ability to check it. No published design gets both.

Frontier Screening cannot see the two things that matter most about an order: intent and assembly. A per-order sequence check is a check on that order. Splitting an order across providers defeats it unless providers can compare detections, which needs either a shared plaintext database nobody will accept or the privacy-preserving log architecture one system has implemented. The published tool limitations are explicit about short fragments, protein-level orders and ambiguity codons, and those limits are architectural rather than budgetary.

Established Institutional capacity is a bottleneck at a scale that makes the policy debate slightly unreal. A 2025 analysis of implementation gaps documents university biosafety offices running multiple compliance mandates with fewer than three full-time professional staff, and most institutions lacking any institution-wide sequence screening capability, commercial screening software, or the ability to audit their own legacy construct inventories. A mandate whose unit of compliance is the institution is executed by those offices.

Frontier Auditing a model at the depth the audit literature says is necessary requires infrastructure that is being built, slowly, by one non-profit and one agency. Secure enclaves, remote execution and attestation certificates exist in demonstration form; their own proponents describe scaling them to frontier-sized models as requiring significant engineering work, and the model owner retains approval rights over which questions may be asked at all.

5 · Research dependencies

Frontier This subject depends on an evaluation science that does not exist yet, and the dependency is methodological rather than biological. What is missing is ordinary measurement theory applied to capability tests: construct validity, a stated relationship between score and the quantity of interest, inter-rater reliability, and an error model. AI Governance establishes this deficit for capability measurement generally; the biological case inherits it and adds a constraint that brief does not face, namely that the criterion variable cannot be observed even in principle without a prohibited experiment.

Established It depends on curation labour that has no stable funder. Every screening tool is a database plus an algorithm, and the database is a hand-curated list of regulated agents, toxins and control-list entries that must track changes in at least three jurisdictions’ export regimes. The open-source tool that publishes its databases does so on an annual-to-quarterly update cadence and has asked for six figures to stay in active development. Curation is the dependency; software is not.

Frontier It depends on standards work that is drafted but not adopted. Two synthesis-quality standards are published in the ISO 20688 series, covering synthesised oligonucleotides in 2020 and gene fragments, genes and genomes in 2024; a third part focused specifically on standardised sequence screening is described as tentative and under consideration. A completed Draft Standard Guide for nucleic acid providers exists to harmonise approaches and make screening data interoperable. Until a screening standard is adopted, “attests to compliance” in a procurement clause has no referent a third party can check.

Speculative It depends on privacy-enhancing computation maturing faster than the access argument hardens. Secure enclaves, multi-party computation and zero-knowledge proofs are the only proposals that let an auditor establish something about a model without the developer handing over weights or the auditor handing over a benchmark. If that tooling matures the access fight becomes negotiable; if not, it stays a question of legal compulsion, and no jurisdiction has yet compelled deep access.

Established It does not depend on any further progress in biological design. Every control described here is buildable against today’s capability. Design progress changes the urgency and the required detection breadth; it unlocks none of the institutional steps, and none of them is blocked by a missing scientific result.

6 · Required experiments

Frontier The decisive experiment is a standing, blinded order-placement audit of the whole provider population, with per-provider results published. It would extend the June 2025 exercise from twelve orders placed once into a continuous regime: benign test orders submitted under varied customer identities at unannounced intervals, scored on whether the order was stopped, queried or shipped, with an adjudication step that distinguishes the three explanations the 2025 exercise could not separate. This is the single result that would most change this brief’s assessment, because it converts the one number everyone quotes — screening sensitivity on a benchmark — into the number that actually governs, which is the probability that a live order is stopped by an organisation. It needs no new science, no new hardware and no new legal theory; it needs a body with a mandate to place the orders and publish the results, and no such body exists or has been funded.

Established The instrumentation for it is already built and running at half scale. The monthly proficiency programme sending 1,000-sequence sets to providers, with 200 graded positives, 200 graded negatives and 600 ungraded sequences, is a proficiency test of tools. Turning it into a proficiency test of providers requires the orders to arrive through the ordinary commercial channel rather than as a labelled test set, and requires the publication of results by provider rather than as a population median. Both are policy choices, not technical ones.

Frontier The second experiment is a published false-negative rate for a deployed biological safeguard stack, produced by a party with white-box access and a fixed adversary budget. The current published record gives categorical descriptions of classifiers and access controls, a bug bounty that found a small number of potentially effective jailbreaks, and an external red team working with content filters disabled. None of that yields a rate. The audit literature specifies what the access would have to be — weights, fine-tuning and internal findings, not query access — and the enclave-and-attestation demonstrations show it can be done without the developer surrendering the weights outright. The obstacle is that no jurisdiction compels it and no developer has volunteered it.

Frontier The third is a head-to-head architecture trial: exact-match against similarity screening, on a common adversarial set generated and held by a third party. The October 2025 disclosure established that one architecture had a problem and that patching improved it; it did not establish a post-patch residual for either architecture, and the two vendors have no incentive to run the comparison themselves. A standards body already runs the benchmark infrastructure this would sit on.

Speculative The experiment everyone wants and nobody can run is a controlled uplift trial with a real criterion variable. The 2024 red-team study is the closest published approach, and its authors reported a null result while stating that the design could not detect the effect it sought and that the fix is more models, more researchers and less variance. Every subsequent proposal has the same structural problem: the outcome that matters cannot be measured without doing the thing the apparatus exists to prevent. Handwave Proposals to resolve this by analogy — treating benchmark performance as a surrogate endpoint of the kind used in clinical trials — assume a validated surrogate relationship nobody has established.

7 · Engineering requirements

Established A working screening layer is an engineering system with four published components, and three of them exist. A curated control-list database with a version identifier; a matching engine fast enough to sit inline with order intake, demonstrated at around 1,100 nucleotides per second per thread; a key and trust architecture, demonstrated as five separated keyservers under a three-of-five threshold with biweekly rotation; and a receipt — a timestamped signature stating which database version an order was screened against. The receipt is the component that turns screening from a promise into an auditable event, and it is the one almost nobody requires.

Frontier Cross-provider detection is an engineering problem that has a demonstrated solution and no adopting institution. Privacy-preserving detection logs already allow the observation that fragments of one flagged construct have been ordered from several providers, without any provider learning another’s customer or sequence data. This is the technical answer to the split-order problem and to the recordkeeping-versus-privacy objection simultaneously. Nothing in the United States, European or United Kingdom instruments requires it or references it.

Frontier Benchtop devices need attestation, not guidance. The engineering ask is pre-market certification of a firmware screening path, tamper-evidence, and a device-side record of what was synthesised — specified the way safety-critical firmware is specified in other regulated equipment. Guidance currently asks manufacturers to integrate screening capability before synthesis occurs; no jurisdiction tests the claim before the device ships.

Frontier On the model side the engineering requirement is evaluation infrastructure, not more classifiers. What is missing is a remote-execution environment in which an external evaluator can run gradient-based and fine-tuning attacks against a frontier model, obtain a rate, and emit an attestation that only the agreed code ran — the shape demonstrated once between a developer and a national institute. Speculative Its proponents say scaling this to the largest models requires significant engineering work, and no published cost estimate exists for doing so at frontier scale.

Established Agentic deployment adds a logging requirement, and it is conventional software engineering. Agent audit logs, tool permissioning, human approval gates and anomaly detection on connected laboratory systems are the practical controls for automated experimentation, and none of them is hard to build. The governance question is not whether they can be built but who is entitled to read the logs, and that is unresolved wherever automated laboratories are discussed, including in Artificial Scientists, which owns the automation capability itself.

8 · Adjacent technologies

Established The closest technical neighbour is the fabrication capability this brief governs. Synthetic Biology records that fabrication became an engineering discipline while design did not, and independently names a post-rescission screening mandate as one of two constraints its own field waits on; this brief is the expansion of that constraint and takes care not to restate its design-failure evidence. Genetic Engineering owns genome writing and editing as a general capability, and is the reason the order-screening chokepoint exists at all: screening governs the purchase of written DNA, so it inherits whatever that field’s cost curve does to order volumes.

Established The closest institutional neighbours divide cleanly. AI Governance owns the general measurement critique of AI statutes, including compute thresholds and the budgets of national evaluation bodies; this brief borrows that finding and contributes the biological half, where an older physical control regime allows a direct comparison of what an instrumented governance layer looks like. Existential Risk Governance owns treaty machinery and the counterfactual problem across prevention bodies; the biological weapons regime and its verification deadlock sit there, and this brief points at them rather than duplicating them.

Frontier Two adjacencies are usually missed. The first is laboratory automation: as ordering, synthesis and experimentation become one continuous workflow, the screening gate and the agent permission system become the same control surface, which is where this brief touches Artificial Scientists. The second is model-weight security, which sits inside the same developer programmes as biological safeguards — egress controls, two-party authorisation, endpoint allowlisting — so a weights breach is both an access-governance failure and a safeguard failure, and no public framework treats it as both.

9 · Institutional requirements

Established The institutional requirement in one sentence: somebody needs the legal power to compel screening, and nobody has it. The United States mechanism was a funding condition, is suspended pending revision, and has no replacement sixteen months after the order requiring one. The United Kingdom published voluntary guidance in October 2024 with a 50-nucleotide expectation, customer verification described as proportionate and practicable, and recordkeeping of flagged orders; it is advisory. In every jurisdiction this brief examined, the operative instrument is a recommendation, and the entity doing the screening is the entity being screened.

Frontier The second requirement is an auditor, and the word is being used loosely by almost everyone. What exists is proficiency testing of tools by a standards agency, voluntary pre-deployment testing agreements between a national institute and model developers, an external evaluation clause in a European code of practice that publishes only summaries, and industry attestation. What does not exist anywhere is a body that can arrive unannounced, place orders, inspect logs, demand a model under evaluation-grade access, and publish an adverse finding over the objection of the party inspected. The audit literature’s own analogy — a nuclear inspectorate with several hundred inspectors and a right of entry — is a statement about what the gap is, and Existential Risk Governance shows the same right-of-access variable separating the verification regime that works from the one that does not.

Established The third requirement is boring and decisive: a funder for shared infrastructure. The benchmark datasets, the curated control-list databases, the free screening tools and the standards drafts are public goods with no revenue model. One of the two free tools has published its funding gap in six figures and a date in November 2026; the other is a foundation-run system quoting five-figure annual costs to providers. A 2025 implementation analysis reached the same conclusion from the institutional side, recommending staffing grants and shared tools as its second recommendation of seven.

Frontier The fourth is a definition, and it is the one a legislature cannot outsource. A December 2025 witness told the House Energy and Commerce oversight subcommittee that there is no comprehensive shared understanding of which biological AI capabilities pose consequential risks, and asked Congress to authorise a standards body to define the narrow subset that warrant oversight, to resource it to track models, agents and autonomous robotics, and to pass legislation governing de novo gene synthesis. Congress has bills rather than statutes: as of an August 2026 legislative summary, at least four measures addressed nucleic acid standards, biosecurity modernisation and engineering biology readiness, plus further bills on AI data standards, cloud laboratories and biological data access. Bills are not coverage.

Speculative The fifth is international, and the only serious candidate is harmonisation through standards rather than treaty. The United Kingdom cost-benefit analysis puts the value of international harmonisation at roughly six times the domestic-only benefit, which is the strongest published argument for pursuing this through an adoptable standard rather than a negotiated instrument — and the standards infrastructure, unlike a treaty, already exists in draft.

10 · Ethical & societal considerations

Frontier Gated access concentrates the power to decide who does biology, and the failure mode showed up on schedule. Vetted-access programmes give trusted partners capability that ordinary users are refused. In August 2026 researchers were abruptly removed from one developer’s restricted-access programme, told their identity could not be verified or that their accounts were ineligible; the company attributed it to a technical error and asked them to reapply, and the affected researchers were reported to be outside the United States and Europe. Whatever the cause in that instance, the structure is the point: eligibility is adjudicated by the vendor, appealed to the vendor, and correlated with geography.

Established Overinclusive screening has a measurable cost borne by people doing ordinary work. The 2025 implementation analysis argues for functional risk tiering precisely because blanket sequence surveillance produces false alarms, unmanaged legacy inventories and compliance load in offices with fewer than three staff, and its first recommendation is to shift away from blanket surveillance rather than to extend it. A governance design that makes routine research slower without a measured reduction in risk is not neutral; it spends real scientific capacity against an unmeasured benefit.

Frontier Recordkeeping is where security and privacy actually trade. A mandate to record orders and sequence data creates an attribution capability and a database of what every laboratory is building. The cryptographic architecture demonstrates that detection across providers can be achieved on hashed data, so the trade is softer than the debate assumes — but only if legislatures write the requirement in terms of a verifiable detection log rather than a plaintext register.

Speculative The deepest ethical question is who is entitled to the window. A 2026 essay describes an embargo window between recognising a capability and its diffusion, and argues that the window is worth having only if it is used for structured pre-release access with clear triggers, independent audits, defined exit criteria and transparent reporting — and that without those it becomes gatekeeping dressed as safety. That is a testable institutional claim, and no current programme meets all four of its conditions.

11 · Civilizational implications

Frontier If the chokepoint thesis is right, this is one of the few places where a small, cheap institution buys a disproportionate amount of civilisational safety. The screening layer costs tens of thousands of dollars per provider per year in published figures, and one analysis puts the return at about three and a half to one. Very little else in catastrophic-risk policy has a quantified ratio at all, and this one does because the control acts on a physical transaction that can be counted.

Frontier If the chokepoint thesis is wrong, the failure is quiet and the institution keeps reporting success. A screening regime measured by benchmark sensitivity and self-attested compliance will continue producing good numbers whether or not it is stopping anything, because neither metric is sensitive to the failure it is meant to prevent. This is the same structural problem Existential Risk Governance documents across prevention bodies: throughput statistics from the body being assessed, and a negative existence claim with an unobservable denominator.

Speculative The durability question is political and the record is poor. A framework issued in April 2024 with a compliance date in April 2025 was ordered revised in May 2025 and had no successor by September 2026. Whatever is built next has to survive administrations, and the only structures in this field that have survived a change of government are the industry consortium, the standards work and the two free tools — none of which can compel anyone.

Speculative The asymmetry cuts both ways and the defensive side is under-argued. The same access architecture that gates capability also enables sponsored defensive work: detection, surveillance, countermeasure design. Handwave The claim that defensive uplift outruns offensive uplift, made in various forms by developers and by some biosecurity groups, is asserted rather than measured, and nobody has proposed a metric on which that race could be scored.

12 · Timelines

These horizons track the governance layer only — what is mandated, measured and auditable — not what models or laboratories can do.

  • 10 yr: Frontier Mandatory screening in at least one major jurisdiction is the most likely single change, with the United Kingdom cost-benefit analysis already recommending legislation by the fourth quarter of 2026 and four United States bills in circulation as of August 2026. A screening-specific ISO standard and a provider guide are drafted and could be adopted inside this window. Whether any of it comes with an inspectorate is the open question; every current instrument is an attestation regime.
  • 25 yr: Speculative Either an audit function with a right of access exists for both halves of the stack, or the physical chokepoint has decayed far enough — benchtop devices, distributed and offshore capacity — that order screening is no longer the control that matters and governance has moved to devices and laboratories. These are the two coherent end states; the current trajectory does not distinguish between them.
  • 50 yr: Speculative If evaluation science matures into something with validity statistics and error models, capability governance becomes a regulatory discipline resembling pharmacovigilance, with post-market surveillance and mandatory incident reporting doing most of the work that pre-deployment testing is asked to do today. If it does not mature, the compute and capability thresholds now in law will have been replaced several times without ever being validated.
  • 100 / 250+ yr: Handwave Any statement at this range assumes the persistence of nation-state licensing of physical manufacture, which is the actual load-bearing assumption of the entire field and has never been examined in this literature. Nothing in the current evidence base supports a claim about it either way.

13 · Technology tree & dependencies

  • Depends on AI Governance for the general finding that capability measurement is the unresolved core of every AI instrument now in force — this brief inherits that result and does not re-derive it, and every claim here about evaluation validity is downstream of it. It also depends on Genetic Engineering, which owns genome writing as a capability and therefore sets the order volumes any screening regime has to process, and on Synthetic Biology, which independently names a post-rescission screening mandate as a constraint its own field is waiting on. Nothing else on this map blocks this topic, and no scientific result is required to build any control described here.
  • Requires (not on this map) A statute rather than a procurement condition: the United States instrument attached screening to federal funding, was ordered revised in May 2025, and had no successor by September 2026, while the United Kingdom and European positions are guidance and absence respectively. A blinded order-placement audit whose per-provider results are published, which is the only way the benchmark medians reported by tools become a statement about what live commerce does; the one published live test placed twelve orders and could not separate three explanations for the three that went through. An evaluator access right in law, because the audit literature shows black-box querying cannot establish the absence of a dormant capability, and every present arrangement is voluntary and publishes summaries. Sustained funding for the shared tools and curated control-list databases, one of which has published a six-figure annual gap and a November 2026 date for entering maintenance mode. Certified benchtop firmware, since current guidance asks manufacturers to build screening in and no jurisdiction tests the claim before shipment. And a validated link between benchmark performance and operational uplift, which is the one item on this list that is a research result rather than a decision, and which the 2026 international assessment still names as an open gap.
  • Enables Everything downstream that depends on biological manufacture remaining a permitted activity. A screening and evaluation regime that is credible is also the political precondition for open scientific access to powerful biological models; the alternative to auditable governance is not freedom but discretionary gatekeeping by whoever holds the weights.
  • Adjacent Synthetic Biology and Genetic Engineering on the fabrication side; AI Governance on evaluation and statutory design; Existential Risk Governance on treaty machinery, verification access and the counterfactual problem; Artificial Scientists where ordering, synthesis and automated experimentation converge into a single control surface.

14 · Common misconceptions & speculative claims

Frontier “AI has already given novices the ability to build biological weapons.” The evidence is genuinely mixed and the strongest claim on each side is worth stating precisely. In favour: a frontier model scored 43.8 per cent on a 322-question expert virology instrument against a 22.1 per cent in-speciality expert baseline, and two developers have judged that they cannot rule out novice uplift. Against: the only published controlled red-team study found no statistically significant difference in plan viability with or without model assistance, and the 2026 international assessment says it remains difficult to judge how far material barriers still constrain actors. The open letter demanding mandatory screening concedes the point in its own text; anyone asserting the claim without that concession is overclaiming.

Established “Screening is solved — the tools score above 95 per cent.” That figure is tool sensitivity on a curated benchmark of 200-base-pair fragments, and it is a real achievement. It is not a statement about whether a live order is stopped. The only published live test placed twelve orders, saw three processed without follow-up, and could not distinguish absent screening from screening that cleared the customer. Until the second measurement exists, the first one cannot carry the argument.

Established “The United States has a synthesis screening mandate.” It does not, and it never had one in the regulatory sense. The 2024 Framework worked through federal funding recipients’ purchasing, was ordered revised or replaced within 90 days in May 2025, and as of September 2026 the operative document is 2023 guidance that is advisory. The July 2026 policy that did issue governs high-risk life-sciences research, not synthesis screening.

Frontier “Generative protein design has permanently broken sequence screening.” The October 2025 Science study did show that screening tools of the time failed to reliably detect variants produced by open-source design software — and the same study coordinated patches with the tool maintainers and reported detection greatly improved. The correct reading is that screening is patchable and was patched. The correct caveat is that no post-patch residual false-negative rate has been published, so “fixed” is as unsupported as “broken”.

Established “Voluntary industry standards are close enough to regulation.” The main consortium is estimated to cover about 80 per cent of global commercial capacity; more than 30 providers are described as voluntarily compliant; more than 150 researchers have signed a biodesign community statement undertaking to buy only from screening providers. None of these arrangements includes an audit, a penalty, or any mechanism that reaches the fifth of capacity outside them.

Frontier “National AI safety institutes are independently verifying biological safety.” They are doing real work: pre-deployment access including reduced-refusal checkpoints, and a reported finding of no broad increase in national-security-relevant biological capability for the most recent flagship model. But the agreements are voluntary, the published outputs to date concern competitor models and cyber and agent security, and no biological evaluation report from a national institute appears on the public record this brief could obtain. Access is not verification, and verification is not publication.

Speculative “Mandatory recordkeeping means a surveillance database of all biological research.” It could, and that is the design most often imagined. It need not: a deployed system already demonstrates cross-provider detection of split orders using privacy-preserving logs in which no party holds another’s plaintext, with signed certificates recording which database version screened an order. The objection is sound against the naive design and weak against the demonstrated one; the risk is that legislation is drafted before anyone notices the difference.

Handwave “Publishing the gaps is how the field gets fixed.” This is the norm imported from software security, and it is the one this brief declines to follow in full. The October 2025 disclosure worked because the vulnerability went to maintainers, was patched, and only then described, with the reproduction detail omitted. A governance brief that publishes a gap map inverts that sequence. Where this brief says a residual is unmeasured, that is a request for the measurement, not an invitation to demonstrate it.