In 2001 the first draft human genome cost roughly three billion dollars and took over a decade. Today a whole genome can be sequenced for a few hundred dollars in about a day. That is one of the steepest cost declines in the history of technology — faster than Moore's law over the same period. If reading the genome were the bottleneck to understanding human biology, we would have understood it years ago. We have not, and this module is about why.
Reading is genuinely solved
Established Modern sequencing works by reading enormous numbers of short DNA fragments in parallel and reassembling them computationally. It is fast, cheap, and accurate. In 2022 the Telomere-to-Telomere consortium closed the last gaps the 2001 draft had left — the repetitive, hard-to-assemble ~8% — producing the first truly complete, gapless human genome. The engineering problem of turning a physical DNA molecule into a string of letters is, for practical purposes, done.
So we can hand you the complete text of your genome. The question this module turns on is: what can we actually tell you from it?
The interpretation gap
Frontier Compared to a reference, any individual genome differs at roughly four to five million positions. Each of those differences is a variant. The clinical question for every variant is the same: does this matter? For the overwhelming majority, we do not know.
Several facts stack up to make this hard. Only about 1-2% of the genome codes for protein; the rest, once dismissed as “junk,” includes vast stretches of regulatory sequence that controls when and how much genes are expressed (the ENCODE project's central finding). Most variants associated with common diseases sit in exactly this non-coding regulatory DNA (as catalogued across thousands of genome-wide association studies), where our ability to predict the effect of a change is weak. And even within coding regions, a change that swaps one amino acid for another may be catastrophic, harmless, or harmful only in combination with other variants or environments.
The formal name for the residue is the variant of uncertain significance (VUS). It is not a rare edge case. Clinical labs generate them by the million. A VUS is the genome saying: here is a difference, and neither you nor anyone else currently knows what it does. Cheap sequencing did not shrink this pile — it made it bigger, by finding more variants faster than we can interpret them.
The genotype-to-phenotype map
Frontier The real object of desire is the genotype-to-phenotype map: the function that takes a genome (plus environment) and predicts the organism. For a handful of conditions this map is nearly complete — single-gene, high-penetrance disorders like sickle-cell (the exact glutamate-to-valine change from M-Bio-04) or Huntington's, where one variant reliably predicts one outcome.
But most traits and common diseases are polygenic: the sum of thousands of variants each nudging risk by a fraction of a percent, interacting with each other and with diet, exposure, and chance. Polygenic risk scores can now stratify populations statistically, but they predict a distribution, not an individual's fate, and they generalise poorly across ancestries because the training data is skewed. The map from sequence to person is, for most of what we care about, still mostly blank.
Where machine learning helps — and where it stops
Frontier The most active frontier is using machine learning to fill in the map. Models like AlphaMissense predict, across the whole proteome, which single-amino-acid changes are likely to be pathogenic, and reclassify large numbers of variants that were previously uncertain. Sequence-to-expression models like Enformer predict how regulatory DNA controls gene expression, reaching further into the non-coding genome than anything before.
These are real advances and they matter. But they share a ceiling, and it is the same ceiling that dominates the AIHS study's hardest bottleneck: they are trained on correlations in observational data, and a prediction that a variant “looks pathogenic” is not the same as knowing what will happen if you intervene to change it in a particular patient. Predicting effect and predicting the consequence of an intervention are different problems — the distinction the causal-modelling bottleneck (advance B3) in the feasibility study is built around.
Why this is the frontier the AIHS study inherits
Speculative Any autonomous healing system must, at minimum, read a patient's molecular state and know what it means. This module is where you learn that the second half of that sentence is the unsolved part. Sequencing hardware is not the obstacle; the obstacle is the interpretation layer — turning a complete molecular readout into an accurate, individual, actionable understanding.
That is not a reason for despair, and it is worth stating in the study's own partial-system framing: you do not need the complete genotype-to-phenotype map to deliver enormous value. Reading the fraction of the genome we do understand already transforms diagnosis for single-gene disorders, pharmacogenomics, and cancer. The frontier is the rest — and being honest about where the understood part ends is exactly the discipline the four-flag system enforces.
Reading a person's genome now costs a few hundred dollars and takes a day. Why does that not translate into being able to tell them what most of their genome means for their health?
Show answer
Because sequencing produces the text, and understanding requires the genotype-to-phenotype map, which is largely unknown. Any given genome contains millions of variants; only a small fraction fall in protein-coding regions, most disease-relevant variants sit in poorly-understood non-coding regulatory DNA, and even for coding variants we frequently cannot say whether a change is harmful, harmless, or context-dependent. The result is millions of “variants of uncertain significance.” Cheap sequencing solved the reading problem; the interpretation problem — connecting sequence to consequence — is a different and far larger scientific challenge.
A funnel widget would work well here: start with the ~4-5 million variants in one person's genome, then apply filters (coding vs non-coding, known vs novel, predicted effect) and watch the count fall — and watch the large residue that lands in “uncertain significance” refuse to disappear. It would make the interpretation gap viscerally countable.