← Blog

The Next AlphaFold Won’t Design Drugs. It Will Find Targets

Evolution has already generated the data. The next challenge is learning what it means.

Title card reading “From DNA to targets”, beside a three-step GI inference diagram: input DNA, a GI genomic world model, and predicted phenotypes and targets as output.

Daphne Koller’s “Drug Discovery Has No Magic Wands” makes an important point: AI in drug discovery is heavily focused on turning known mechanisms into drugs, especially by designing molecules. But many programs fail much earlier, when researchers move from disease biology to a mechanism and choose the wrong target.

That imbalance may define the next major opportunity for AI in biology.

Chart titled “Where AI Effort Goes vs. Where Drugs Fail” across three stages of drug discovery. AI effort is Low at disease→mechanism, High at mechanism→drug, Moderate at drug→patient; failure origin is the mirror image — High, Low, Moderate. Caption: the effort is concentrated where the problem isn’t.

Where AI effort concentrates versus where programs actually fail. Illustrative; insitro / a16z.

AlphaFold showed what happens when the right representation, enough data, and enough compute meet an important biological problem. It triggered an enormous wave of work around protein structure, protein interactions, molecular generation, and therapeutic design. But all of these become much less useful if we are designing an excellent molecule for the wrong biological mechanism.

The next AlphaFold-scale breakthrough may happen upstream, with a model that discovers targets.

We are getting much better at molecules than targets

AlphaFold was a scientific milestone because it demonstrated that evolutionary sequence information could be converted into remarkably accurate predictions of protein structure. AlphaFold2 explicitly uses homologous sequences and evolutionary relationships as key inputs. (Nature)

Its success naturally attracted capital and talent toward adjacent problems: protein design, binding, molecular interactions, and generative chemistry. A recent a16z discussion of biological foundation models, for example, focuses heavily on molecular design and biomolecular interaction modelling. (a16z)

The work matters, but it begins after someone has answered the harder question:

What biological mechanism should we intervene on?

Target discovery is still an evidence-integration problem

There is no single target-discovery pipeline. Targets can emerge from human genetics, GWAS, transcriptomics, proteomics, patient data, CRISPR screens, phenotypic assays, known pathways, clinical observations, and published research. Platforms such as Open Targets integrate many of these evidence sources to prioritize target–disease associations.

In simplified form:

Genetics + omics + phenotypes + perturbations + literature

Candidate prioritization

Experimental validation

Therapeutic hypothesis

This is far better than reading papers and guessing. It is also expensive and fragmented, and it depends heavily on the biology we have already measured. We see a small part of an enormous biological state space, combine the available evidence, form a hypothesis, run an experiment, and repeat.

The natural response is: generate much more biological data.

Maybe.

There may be another route to much of that missing representation.

A lesson from language models

Here is a strange fact about large language models.

An LLM has never seen a chair.

It has never watched a company collapse, crossed a street, raised a child, or performed an experiment. Most LLMs were trained primarily on language produced by human brains.

And yet language carries a compressed trace of the world people have observed. By learning enough structure from that trace, LLMs acquire representations that support reasoning about economics, physics, psychology, software, history, and much else. They do not have to reconstruct the photons that hit the retina, simulate every neuron firing in the human brain, and then regenerate the sentence. Those intermediate levels are implicit in the representation.

Biology may work similarly.

Today, an LLM can work with what humans already know about biology because we have written that knowledge down. But it cannot look directly across billions of genomes and discover regularities that no one has yet described.

Sequence is a different modality. It needs its own foundation models. And those models may not need to simulate every intermediate biological layer explicitly. A sufficiently powerful genomic language model could learn a latent representation in which information about proteins, regulation, cellular function, organismal fitness and, ultimately, phenotype is already embedded.

A genomic foundation model might then be aligned with phenotype using relatively small amounts of carefully chosen experimental and clinical data, much as an LLM can be adapted to a specific task with limited task-specific data.

That is the hypothesis.

Rather than relying on a long, complicated pipeline:

Genome → model every molecule → model every cell → model every tissue → predict phenotype

we may be able to use something shorter and more scalable:

Genome → learned genomic world model → phenotype

The intermediate biology still exists, of course. The model simply may not need to represent every layer explicitly.

Evolution has already generated the training signal

Where would such a representation come from?

Evolution.

Mutation changed sequences. Recombination generated combinations. Selection constrained which configurations survived and reproduced. Co-evolution accumulated information about which biological components could function together.

Genomes are not merely strings. They are the surviving output of billions of years of biological experimentation.

This is not equivalent to having randomized perturbation experiments with perfectly measured phenotypes. Environmental context is often missing from in vitro systems. Lethal variants disappear. Late-onset human disease is only weakly coupled to reproductive fitness. But the evolutionary record still contains an extraordinary amount of information about biological constraints.

We already have sequence at an extraordinary scale

Two panels. Left, “The Data Chasm” on a log scale: single-cell 3–25 GB, the AlphaFold corpus ~1.2 TB, a cell atlas 10–50 TB, a cell-type dictionary 10–100 TB, against the 10–100 PB a causal map of biology would require — a gap of about three orders of magnitude. Right, database size in petabases: BLAST nt 0.001, BLAST WGS 0.024, raw SRA 2024 50.0, Logan unitigs 4.6, Logan contigs 0.9.

The gap between the data we have and the data a causal model of human biology would need — and the sequence scale Logan makes searchable. Illustrative; insitro / a16z, with Logan figures from Chikhi et al.

The raw sequence-data problem is becoming very different from even a few years ago.

The Logan project processed roughly 50 petabases from more than 27 million public sequencing datasets, turning the Sequence Read Archive into a searchable resource at planetary scale. (Logan paper)

This points to a different strategy than mapping biology experimentally from scratch.

Pretrain first on evolution. Then spend expensive experimental data on alignment.

The genomic world model

Imagine a genomic foundation model pretrained across the diversity of life, then aligned with:

  • human genetics
  • clinical phenotypes
  • expression and epigenetics
  • single-cell measurements
  • perturbational experiments
  • molecular and structural data

A good model should eventually be able to address questions such as:

  • What phenotype is likely to result from this perturbation?
  • Which genes are upstream drivers rather than downstream consequences?
  • Which targets could alter disease without creating unacceptable liabilities?
  • Which experiment would most efficiently distinguish between competing mechanisms?

If neural scaling laws carry over, the model should generalize to biological combinations that have never been measured directly. At that point it starts to look like a genomic world model, not another database.

Experiments become alignment

This changes the role of experimental data. Today, much biological knowledge has to be constructed through expensive measurement, one context at a time. In a foundation-model regime, much of the representation could be learned before the first disease-specific experiment.

Experiments would then provide the scarce information that evolution cannot: human-specific phenotypes, intervention effects, disease context, and clinical outcomes. Natural-language processing offers a useful precedent. Transfer learning can sharply reduce the amount of task-specific downstream data needed to adapt a model.

The new workflow is:

Evolutionary pretraining
→ phenotype alignment
→ target prediction
→ high-information experiment
→ model update

With a sufficiently rich pretrained representation, relatively modest amounts of disease-specific data might unlock predictions across several biological levels without simulating every intermediate mechanism. That could change the economics of target discovery.

The next bottleneck

For the last few years, the obvious AI-biology question has been:

Can AI design better molecules?

Increasingly, the answer is yes.

Now the question is:

Can AI tell us which biology is worth targeting?

Evolution has already generated an enormous training corpus. Genomic foundation models are beginning to learn its language.

The key question is how much phenotype is latent inside that language. How little additional data would we need to extract it?

If the answer is “a lot” and “not much,” the next AlphaFold-scale company will not generate molecules. It will tell everyone else which targets to build them for.