1 · Concept overview
Established Generative biomolecular design is the use of learned models to propose molecules — proteins first, nucleic acids and their assemblies increasingly — that do not exist in nature and are then made and tested. The field has a Nobel Prize, a dozen named methods, and one measured quantity that decides whether any of it is engineering: the fraction of designs that work when they reach a bench. Frontier This brief is organised around that quantity and around the four things that stand between a binding design and a product — function, immunogenicity, manufacturability and independent replication.
Established The best aggregate number in the published literature is one in nine. A 2025 meta-analysis assembled 3,766 experimentally characterised de novo binders across 15 targets and found 436 — 11.6 per cent — bound in vitro, with per-target success ranging from 0.1 to 1.0 and the authors stating that design success remains very inconsistent and that there are no standard criteria for prioritising which designs to test. Frontier It is a preprint, it is the only aggregate figure available, and it is carried at that weight here. Established The Synthetic Biology brief uses the same number as one of two measurements of prediction failure; this brief takes it as the entry point to a different question — what happens to the nine designs in ten that did not bind, and to the one that did.
Established The clinical state can be stated in one sentence: no de novo designed protein binder is approved as a medicine anywhere. Established The nearest thing to a counterexample is a computationally designed self-assembling nanoparticle used as a vaccine scaffold, approved in South Korea in 2022, in which the designed part is the scaffold and the immunogen it displays is natural. Frontier The first de novo designed protein therapeutic to enter human trials, an engineered cytokine mimetic, was discontinued by its sponsor in 2023. Frontier Those two facts, held together, are the honest summary of where translation stands.
Frontier The gap this brief is most interested in is between benchmark and bench. Design campaigns are filtered in silico before anything is synthesised, and the filters are structure-prediction confidence scores whose prospective predictive value is precisely what the meta-analysis above set out to measure. Established Method papers report the targets their authors chose; independent laboratories testing community-submitted designs report lower rates. Speculative Whether the difference is target selection, reporting practice or method quality is not currently separable from the published record, and saying so is more useful than picking one.
Established A note on sourcing. This brief was commissioned in September 2026 from the Institute’s research base. Reading-list entries without links are cited from the bibliographic record rather than re-fetched, and claims are dated no later than early 2026 unless carried by a linked source.
2 · Current scientific position
Established The 2024 Nobel Prize in Chemistry is the field’s institutional arrival and is routinely misread. It was awarded in two halves: one to David Baker for computational protein design, the other jointly to Demis Hassabis and John Jumper for protein structure prediction. Established Those are different problems. Prediction takes a sequence and returns a structure; design takes a desired structure or function and returns a sequence, and the reason prediction matters to design is that it supplies the cheap filter which decides what gets synthesised.
Established Prospective success rates, as published, are the spine of any honest assessment, and they are reported in incompatible units. The aggregate figure is 11.6 per cent of 3,766 tested designs across 15 targets, with per-target success spanning the full range from 0.1 to 1.0. Frontier Method papers report their own campaigns: diffusion-based backbone generation reports per-target hit rates from well under one per cent to tens of per cent depending on the target; the BindCraft pipeline reports experimental success between roughly 10 and 100 per cent on a small panel its authors selected; and a commercial laboratory’s AlphaProteo system reported binding rates of roughly 9 to 88 per cent across seven targets in a technical report rather than a peer-reviewed paper (the developer’s own figures). Frontier None of these numbers is directly comparable to the others, because the denominator is chosen by the designer: how many designs were made, how they were filtered, and how many were carried to the assay are all discretionary and rarely fixed in advance.
Established What counts as success in nearly every one of these reports is in-vitro binding, not function. A design succeeds if it is expressed, is soluble, and shows measurable affinity for its target by yeast display, biolayer interferometry or surface plasmon resonance. Frontier Cellular activity, specificity against a proteome, stability in serum and activity in an animal are separate hurdles that most published campaigns never attempt. Established A one-in-nine binding rate therefore is not a one-in-nine rate of useful molecules, and the field’s own reviews say so.
Frontier Beyond binding, the strongest published demonstrations of designed function are narrow and real. A protein language model produced a fluorescent protein sharing roughly 58 per cent sequence identity with the nearest known natural fluorescent protein and it fluoresced, which is a genuine novel function rather than a re-derivation. Frontier De novo enzymes remain the hardest case: the best published designed hydrolases reach catalytic efficiencies orders of magnitude below their natural counterparts, and improvement still comes largely from laboratory evolution applied after the design step. Established Computational redesign has nonetheless delivered at least one deployed capability — the redesigned essential enzymes behind synthetic auxotrophy for biocontainment, which reported escape frequencies below the limit of detection.
Frontier Agent-orchestrated design is now in the same evidentiary frame, and its best result is modest and clearly reported. An agent-directed pipeline combining a language model, structure prediction and physics-based scoring designed 92 nanobodies against SARS-CoV-2 variants, of which two showed improved binding to the variants of interest. Established That is a publishable result and a small one, and the Artificial Scientists brief carries the general lesson it illustrates: where a cheap exact verifier exists the loop closes, and where the verifier is a wet-lab assay the throughput of the assay sets the pace.
Established The one place designed proteins have reached a regulator is vaccine scaffolds. A two-component self-assembling nanoparticle, computationally designed, displaying a natural coronavirus antigen, was approved in South Korea in 2022. Frontier The mosaic nanoparticle line of work is the research frontier of the same idea: a designed 60-subunit particle displaying receptor-binding domains from multiple sarbecoviruses, with published rabbit immunogenicity across a twelve-pseudovirus panel and macaque protection against heterologous challenge reported in the peer-reviewed literature. Established The Universal Vaccines brief owns what happens next to candidates like these, and its finding is the discipline this field needs: until an immune measurement predicts protection, encouraging early-phase immunogenicity is uninformative about efficacy.
Frontier Antibodies are the commercially dominant format and the hardest published case. De novo antibody design has to place a binding surface on loops whose conformations are themselves hard to predict, and the reported campaigns return low single-digit hit rates with affinities in the micromolar range that require experimental optimisation before they are useful. Established Every approved antibody drug to date came from an immunised animal, a display library or a human donor, and none came from a generative model. Frontier This is the cleanest illustration of a pattern that runs through the whole subject: success rates fall as the target format moves from a rigid designed interface towards the formats medicine actually uses.
Frontier Immunogenicity is the least-evidenced link in the entire chain, and the field mostly reasons about it by analogy. A de novo protein is by construction a sequence the human immune system has never seen, and anti-drug antibodies are the standard failure mode of protein therapeutics, with incidence across approved biologics ranging from negligible to the large majority of treated patients depending on the product. Frontier There is no published clinical anti-drug-antibody dataset for a de novo designed binder, because there has been no sustained clinical programme to generate one. Speculative The argument that small hyperstable designs should be less immunogenic than antibodies — fewer epitopes, faster clearance, no Fc — is plausible and untested; the opposing argument, that novel folds present novel T-cell epitopes, is equally plausible and equally untested. Handwave Any claim that de novo proteins are inherently low-immunogenicity is currently an assertion.
Established Manufacturability constraints are ordinary, well understood and routinely omitted from design papers. Expression yield, solubility, aggregation and purification behaviour decide whether a design can be made at scale, and designs are typically screened for them after the fact rather than selected for them. Established Pharmacokinetics is the sharper constraint: proteins below roughly 60 kilodaltons are cleared renally in minutes to hours, so a designed miniprotein that binds beautifully needs a fusion partner, an albumin binder or conjugation before it can be dosed, and each of those changes the molecule that was designed. Frontier Nothing in the current design objective function captures this, which is why the molecule that reaches manufacture is rarely the molecule the model proposed.
Frontier Independent replication is the field’s weakest institution. The measured evidence that third-party testing gives lower hit rates than method papers comes from open design competitions run by commercial testing laboratories, in which community-submitted designs against a common target were expressed and assayed under one protocol and most did not bind. Frontier This brief cites those competitions by name from the bibliographic record and prints no round-level numbers, because it could not verify them. Established What can be stated from the verified literature is the meta-analysis’s own conclusion: success is very inconsistent across targets and there are no standard criteria for choosing which designs to test — which is a statement about the field’s reporting practice as much as about its methods.
Frontier What “de novo” means in this literature is narrower than the phrase suggests. It means a sequence not derived from a natural homologue, not a structure unlike anything in nature: the great majority of designs occupy folds already represented in the structural databases, which is unsurprising because those folds are what the models were trained on and what prediction scores them against. Established Novelty is measurable — sequence identity to the nearest natural relative, structural similarity to the nearest known fold — and the designs most often cited as genuinely novel are those whose authors published that distance. Speculative Whether generative models can reach regions of structure space that evolution did not sample is an open question the published record does not settle.
Frontier Generative design has moved past proteins, and the capability claims there are earlier and looser. Genome-scale language models trained on nucleotide sequence now generate candidate regulatory elements, genes and, in 2025, whole bacteriophage genomes reported as viable in culture. Speculative Assessed strictly as capability, this is a change of scale rather than of kind: the verification problem grows with the size of the object, and a generated genome has no cheap exact verifier at all. Established The misuse and access-control questions this raises belong to the AI-Biology Governance brief, which assesses the screening stack those capabilities run against; this brief does not duplicate that assessment and takes no position on access policy.
3 · Frontier questions
Frontier Is the hit rate improving, and by how much? The field has no standing metric, because each paper chooses its own targets, filters and denominators. Frontier A time series of first-pass success on a fixed target panel would answer in a year what the current literature cannot answer at all, and the meta-analysis that produced the 11.6 per cent figure exists precisely because no such series does.
Frontier Are designs novel, or recombinations wearing new sequences? Designed proteins are routinely novel at the sequence level while occupying folds that already appear in the structural databases, and the fluorescent-protein result is notable exactly because its distance from the nearest natural relative was measured and stated. Speculative Novelty claims without a stated structural and sequence distance are not assessable.
Frontier Do the methods generalise to the targets that matter clinically? Published successes cluster on well-ordered, rigid, non-glycosylated epitopes. Frontier Membrane proteins, heavily glycosylated surfaces, intrinsically disordered regions and conformationally mobile targets are where most unmet medical need sits and where per-target success rates fall towards the bottom of the published range.
Speculative Can design replace affinity maturation rather than feed it? Most reported high-affinity designed binders reached their final affinity through experimental optimisation after the generative step. Frontier A campaign that produced picomolar affinity and demonstrated specificity with no post-hoc optimisation would be a categorical change, and none has been published.
4 · Technological bottlenecks
Established Wet-lab throughput, not model quality, sets the rate of progress. The number of designs a laboratory can express, purify and assay per target per month is small compared with the number a model can propose per hour, so every campaign is a filtering problem whose filters are themselves models. Frontier This is why the prospective value of in-silico confidence scores is the central methodological question rather than a technical detail.
Established Negative results are not published, and the aggregate rate is therefore an upper bound. Campaigns that yielded nothing rarely reach print, and the meta-analysis could only pool what had been reported. Frontier A field whose denominator is discretionary and whose failures are invisible cannot compute its own success rate, which is a reporting bottleneck with no technical component.
Frontier Specificity is measured far less often than affinity. A binder that engages its target at nanomolar affinity and also engages three unrelated human proteins is a failure that a standard binding assay will not reveal. Established Proteome-wide counter-screening is expensive, rarely performed in design papers, and is a routine requirement in drug development.
Frontier Structure-dependence limits the target list. The dominant design approaches need a structural model of the target surface to design against, so targets without a reliable structure are effectively out of scope. Frontier Prediction has widened that list considerably and has not removed the constraint for flexible or context-dependent epitopes.
Established The molecule that is designed is not the molecule that is dosed. Half-life extension, formulation, stability and delivery modify the designed entity, and none of those modifications is part of the design objective. Frontier Bringing developability into the objective function is an obvious step that the published literature has barely begun.
5 · Research dependencies
Established This subject consumes structure prediction and produces nothing that structure prediction needs. The filters, the scoring and much of the validation are downstream of prediction quality, and a prediction failure mode propagates directly into a design campaign as an unrecognised false positive.
Established It depends on fabrication it does not own. Gene synthesis, assembly and expression are the manufacturing base, and the Synthetic Biology brief records both what that base delivers and what it does not: writing megabases is priced and industrialised, while predicting what a written sequence will do is not.
Frontier It depends on immunology it cannot currently obtain. A predictive model of clinical immunogenicity for novel sequences would change the risk profile of every therapeutic programme in the field, and building one requires clinical anti-drug-antibody data from de novo proteins, which requires trials that have not been run.
Frontier And it depends on independent characterisation capacity. The claim that a method works is a claim about what happens in someone else’s laboratory, which makes contract and cloud testing capacity part of the epistemic infrastructure rather than a convenience.
6 · Required experiments
Established The decisive test is a standing, pre-registered, third-party prospective benchmark: a fixed panel of targets spanning easy and hard classes, a fixed number of designs per method, one independent laboratory, one assay protocol, and publication of every result including the nulls. It is decisive because every contested quantity in this brief — whether hit rates are improving, whether method papers are target-selected, whether in-silico filters have prospective value — is a direct readout of that design and is unobtainable without it. Frontier Pilot versions already exist as open design competitions run by commercial testing laboratories, which is what makes the standing version a funding decision rather than a research problem. Established Nobody has funded the standing version.
Frontier Second: a first-in-human trial of a de novo designed binder with immunogenicity as a published endpoint. The immunogenicity question cannot be settled preclinically, and a single well-reported phase 1 with anti-drug-antibody incidence, titres and neutralising fractions would constrain the field more than a hundred additional in-vitro campaigns. Frontier The one de novo protein therapeutic that entered humans was discontinued before such a dataset accumulated in the public record.
Frontier Third: a specificity audit at proteome scale on a published design set. Take designs already reported as successes, counter-screen them broadly, and report how many are selective. Established This experiment needs no new methods and no new designs, which is why its absence is a statement about incentives.
Speculative Fourth: a designed enzyme that reaches natural catalytic efficiency without laboratory evolution. This is the cleanest single demonstration that design has become derivation rather than search. Frontier The published record does not contain it, and the gap between designed and natural catalytic efficiency is the most durable quantitative deficit in the subject.
7 · Engineering requirements
Established The engineering requirement that matters most is closed-loop test capacity. Automated expression, purification and binding characterisation at hundreds of designs per target per week turns design from a publication activity into a measurement activity, and the constraint is capital equipment and protocol standardisation rather than novelty.
Frontier Assay standardisation is the unglamorous prerequisite for every comparison in this brief. Yeast display, biolayer interferometry and surface plasmon resonance do not return the same answer for the same molecule, and until a common reporting standard exists, cross-paper hit-rate comparison is approximate by construction.
Frontier Developability screening belongs early, not late. Aggregation propensity, expression yield, thermal stability and serum stability can all be measured cheaply on small sets, and promoting them from post-hoc characterisation to selection criteria is the most obvious available improvement in campaign design.
Speculative Compute is not the binding constraint and is often presented as one. Generating candidate designs is cheap relative to testing them; the expensive resource is the bench, and an order of magnitude more compute buys proposals nobody can assay.
8 · Adjacent technologies
Established Synthetic Biology owns the fabrication half and the diagnosis this brief inherits. Its verdict is that writing is an engineering discipline and design is not, with 11.6 per cent of binders binding cited alongside 19 of 88 designed genome segments proving lethal. Frontier This brief accepts that diagnosis and extends it in one direction only: past the binding assay, into function, immunogenicity, manufacture and replication, where the numbers are thinner and the consequences larger.
Established Artificial Scientists supplies the verification frame. Its central finding is that autonomous discovery held up where a cheap exact verifier existed and failed where the verifier was a noisy instrument. Frontier Biomolecular design is the case where the verifier is a wet-lab assay: neither cheap nor exact, which predicts precisely the reporting pathology this brief documents.
Established Universal Vaccines supplies the translational discipline and AI-Biology Governance the misuse assessment. The first records what happens when immunogenicity is encouraging and efficacy is not measured; the second assesses screening, access control and capability evaluation for exactly the tools described here. Established This brief refers to both and duplicates neither: it makes no claim about access policy and no claim about what a clinical endpoint should be.
9 · Institutional requirements
Frontier The field’s reporting norms are its main institutional deficit. Hit rates are reported per paper, per target, on denominators the authors choose, and negative campaigns are largely unpublished. Established A standing disclosure convention — designs proposed, designs tested, designs that bound, assay used — would cost nothing and is the single change that would make the strong capability claim testable.
Established Regulators have no category for a computationally designed protein, and do not need one. A designed binder would be evaluated as a biologic on the same evidence any other biologic needs: potency, purity, stability, pharmacokinetics, immunogenicity and clinical benefit. Frontier The consequence is that computational novelty confers no regulatory advantage, and the design step compresses discovery timelines without touching the part of the timeline that dominates.
Frontier The capital structure is unusual and matters for what gets published. Much of the method development sits in a small number of academic institutes and well-funded private laboratories, some of which publish technical reports rather than peer-reviewed papers and choose their own evaluation targets. Established Where a developer reports its own success rates, this brief labels it as such, because the labelling is the only correction available to a reader.
10 · Ethical & societal considerations
Frontier The nearest ethical hazard is expectation, not misuse. A Nobel Prize and a stream of striking figures have produced public expectation of designed cures on a timescale the clinical record does not support, and patients hear capability claims as availability claims. Established No de novo binder is approved; the responsible statement of the field’s state says so first and describes the promise second.
Established Misuse and access control are real and are assessed next door. The AI-Biology Governance brief holds the screening and evaluation evidence, including the demonstration that generative redesign can evade sequence screening in silico and the disclosure process that followed. Frontier This brief records only the capability-assessment boundary: a field that cannot predict whether its designs will bind is also a field whose risk assessments cannot rest on predicted function alone.
11 · Civilizational implications
Speculative If prospective success rates rose to the point where designing a functional binder were routine, the scarce input in biologics would shift from discovery to trials. Discovery is a minor share of the cost and time of bringing a biologic to patients, so a field that solved design completely would compress the part of the pipeline that is already cheapest. Frontier The larger civilizational effect would land outside medicine: enzymes for industrial process chemistry, binders as reagents and diagnostics, and designed materials, where there is no clinical trial between a molecule and its use.
Frontier The durable asset is a characterised negative record. Millions of designs tested and reported — including the failures — would be the training data that makes the next generation of models predictive rather than generative, and it is currently being discarded. Speculative A field that published its nulls would be building the dataset it needs; a field that publishes only its hits is training its successors on its own survivorship.
12 · Timelines
These horizons track prospective wet-lab performance and translation, not model releases.
- 10 yr: Frontier A standing third-party benchmark exists and first-pass binder success on a fixed panel is reported as a time series; designed binders are routine laboratory reagents and diagnostics; at least one de novo designed protein reaches a phase 2 trial with published anti-drug-antibody data; designed enzymes remain below natural catalytic efficiency and continue to be improved by laboratory evolution.
- 25 yr: Speculative A de novo designed binder is approved as a medicine, on the same evidence any biologic requires; developability and immunogenicity enter the design objective rather than the post-hoc screen; per-target success rates for well-structured targets exceed half, with hard target classes still an order of magnitude behind.
- 50 yr: Speculative Design becomes derivation for a defined class of problems — a specified catalytic activity or binding profile returns a molecule that works first time — and the bench becomes a confirmation step rather than the selection step.
- 100 / 250+ yr: Handwave Whole-organism or whole-pathway generative design with predicted phenotype, which requires a sequence-to-phenotype model that no current programme is producing and that the neighbouring briefs identify as the binding constraint of the entire area.
13 · Technology tree & dependencies
- Depends on Synthetic Biology for the fabrication base and for the diagnosis this brief extends — writing industrialised, design not — and Artificial Scientists for the verification frame that explains why a field whose verifier is a wet-lab assay reports the way this one does. Universal Vaccines supplies the translational discipline for the one product class where designed proteins have reached a regulator. What this subject is actually blocked on is not on any map: a prospective, third-party measurement of its own success rate.
- Requires (not on this map) Six things that are not new algorithms. A standing third-party prospective benchmark, without which no claim about improvement over time is assessable. A disclosure convention covering designs proposed, tested and successful, since a discretionary denominator makes any published rate an upper bound. A predictive model of clinical immunogenicity for novel sequences, which requires human data that no completed programme has generated. Developability and expression screening at the throughput designs are produced, so that the molecule selected is the molecule that can be made. Independent characterisation capacity, because a method claim is a claim about someone else’s bench. And capital willing to take a de novo protein into humans, which is a financing decision that the discontinuation of the field’s first clinical candidate made harder.
- Enables Reagents and diagnostics with designed specificity; vaccine scaffolds, where the only approved product containing a designed protein already sits; biocontainment and industrial enzymes, where no clinical trial stands between a molecule and its use; and a measured baseline other fields can be held to — a claim of designed function now has to state its assay, its denominator and its independent replication.
- Adjacent Synthetic Biology, Artificial Scientists, Universal Vaccines and AI-Biology Governance, which owns the misuse and access-control assessment this brief deliberately does not duplicate.
14 · Common misconceptions & speculative claims
Frontier “The Nobel Prize means protein design is solved.” The prize was split between design and structure prediction, and it recognised a body of work whose own best aggregate first-pass success rate is 11.6 per cent. Established The Synthetic Biology brief addresses the related claim that AI has made protein design reliable; the point here is narrower and about the prize itself, which was awarded for opening a capability, not for closing it.
Handwave “Published success rates are around half.” Rates near or above half exist and come from method papers reporting targets their authors selected, on denominators the authors set. Established The only pooled figure across laboratories and targets is one in nine, with per-target values spanning the entire range from 0.1 to 1.0, and any single headline rate quoted without its target and denominator is uninterpretable.
Frontier “A designed binder is a drug.” Binding in vitro is the first of at least six hurdles: specificity, cellular activity, stability, pharmacokinetics, immunogenicity and manufacturability. Established No de novo designed binder has cleared them to approval anywhere, and the first one to enter human trials was discontinued by its sponsor.
Speculative “De novo proteins are less immunogenic because they are small and stable.” There is no clinical dataset behind this. Frontier The mechanistic arguments run in both directions and the field has no human anti-drug-antibody data on de novo binders to adjudicate between them, which makes the claim a hypothesis that a single well-reported phase 1 could test.
Frontier “Structure prediction solved folding, so design follows.” Prediction answers what a sequence does; design asks what sequence would do a specified thing, and the second is not the inverse of the first. Established Prediction contributes to design as a filter, and the prospective value of that filter is the question the field’s own meta-analysis was written to measure.
Frontier “Generative models routinely invent new folds.” Designs are commonly novel in sequence while occupying folds already present in the structural databases. Established A novelty claim is assessable only with a stated distance to the nearest natural relative, which is why the designed fluorescent protein at roughly 58 per cent identity to its closest known natural counterpart is cited so often: the distance was measured.
Handwave “These tools can already design a working bioweapon.” The published demonstration that is usually meant — generative redesign of toxin sequences past biosecurity screening — was in silico only, was disclosed to screening providers and patched. Established The full assessment of that episode, and of what evaluation evidence does and does not exist, belongs to the AI-Biology Governance brief and is not repeated here.