← Blog

Applications of DNA language models: from reading genomes to writing biology

The first wave of DNA language model applications is already here — variant interpretation, target discovery, genome annotation, and sequence design — and where the field could go as the models get more powerful.

  • dna-language-models
  • research
  • perspective
Infographic titled 'Applications of DNA Language Models — From Reading Genomes to Writing Biology,' organized into two rows. The top row, 'The First Wave: Applications Already Here,' shows five tiles: variant interpretation, from variants to drugs, genome annotation and feature discovery, sequence design and engineering, and therapeutic and phage applications. The bottom row, 'The Next Wave: What Becomes Possible,' shows seven tiles: programmable cell-state control, synthetic gene circuits, trait borrowing across species, evolutionary forecasting, simulating artificial evolution, closed-loop design platforms, and DNA computation and storage.
The two waves of DNA language model applications — the first already emerging, the second becoming possible as the models grow more powerful.

Over the past few years, protein biology has been transformed by foundation models. Tools such as AlphaFold are now routinely used in drug development. In the DNA world, the story is still earlier. Today, I want us to think about existing and emerging applications of DNA language models, because I believe their impact may become even broader than the impact of protein modeling tools.

The first wave of applications is already here

Even though DNA foundation models are still young, their application landscape is no longer hypothetical. A first tier of applications has already emerged across recent papers and model releases. These applications do not yet amount to a single AlphaFold moment for DNA, but together they show where the field is moving.

The most immediate area is variant interpretation. Many disease-associated mutations lie outside protein-coding regions, in parts of the genome that regulate when, where, and how genes are expressed. These variants are notoriously difficult to interpret. DNA models offer a way to predict whether a mutation may disrupt a promoter, enhancer, splice site, transcription-factor binding site, chromatin state, or other regulatory feature. Models such as GENA-LM, Evo 2, Nucleotide Transformer, and GPN-MSA all point toward this use case: using sequence alone to estimate the functional consequence of genetic variation. You can use our platform at genomicintelligence.ai to see how this works in action, for example, by predicting how variants change gene expression.

Closely related to this is using variant interpretation to develop new drugs. It is well known that pharma targets supported by genetic evidence — that is, evidence for a causal relationship between a molecular target and disease — have a higher probability of success in clinical trials. To find, or prove, target-disease relationships, we often use genetic associations. Yet associations are not necessarily causal, and therefore we often miss the true target. With a powerful eQTL interpretation model, we can assign genetic variants to their targets and specific contexts. For example, we can infer which gene, and in which cell-type context, is affected by a non-coding variant associated with disease. This would increase the value of currently known genomic associations by transforming them into actionable targets.

The second application is genome annotation and feature discovery. Modern genomes contain many elements that are still poorly annotated, especially in non-model organisms and in non-coding regions. DNA language models can help identify genes, exons, introns, promoters, enhancers, polyadenylation sites, mobile elements, viral sequences, and other functional regions. In this sense, they can become annotation engines. Sequencing is cheap, and we are now waiting for efficient genome interpretation engines that can make use of the enormous amount of sequenced data. To facilitate these applications, we have recently launched genome interpretation tools on our platform.

The next category is sequence design. Once models can predict how DNA sequences behave, they can also be used to design new ones. Early examples already include synthetic regulatory elements with cell-type-specific activity, generated genomic sequences, programmable gene-insertion systems, antimicrobial peptides, and AI-designed bacteriophages. This moves DNA models from interpretation to engineering: not only asking what a sequence means, but also asking what sequence we should write.

A striking example of this direction is Trogenix’s work on synthetic super-enhancers for glioblastoma. Their approach uses engineered DNA regulatory elements that are activated in specific diseased cell states, effectively acting as molecular switches for targeted gene therapy. In preclinical glioblastoma models, Trogenix reported complete tumour elimination in 83% of treated cases after a single dose, with no recurrence over 11 months. This is not a DNA language model application by itself, but it shows exactly the kind of design problem where better DNA models could become powerful: finding regulatory sequences that are active only in the desired disease context.

Another important example is phage design. Bacteriophages are viruses that infect bacteria, and they are increasingly interesting as potential therapies against antibiotic-resistant infections. Recent work with genome language models such as Evo 1 and Evo 2 showed that AI-generated phage genomes can produce viable viruses. In one reported experiment, AI-generated phage cocktails overcame bacterial resistance in strains where the original ΦX174 phage failed. This suggests a future where DNA models could help design phage therapies that adapt faster than bacterial resistance.

Taken together, these applications define the first explored tier of the DNA language model field: variant interpretation, target discovery, genome annotation, synthetic sequence design, therapeutic regulation, and phage engineering. The models are still young, but the use cases are already visible.

What might become possible when DNA language models become truly powerful?

The current generation of DNA language models can already predict regulatory activity, annotate genomes, interpret variants, and design short functional sequences. But these applications may only be the beginning. If DNA models become much more accurate, more causal, more controllable, and more aware of cellular context, the next wave of applications could look very different from today’s.

One possibility is programmable gene regulation and synthetic cell-state control. Today, we can sometimes design promoters or enhancers that are active in a particular cell type. A more powerful DNA model could go further: it could design regulatory programs that activate only in a specific disease state, developmental stage, stress response, or immune context. Instead of a gene therapy vector that is simply “liver-specific” or “neuron-specific,” we could imagine vectors that behave like biological logic gates: active only in glioblastoma-like cells, silent in healthy neural tissue, activated under hypoxia, or tuned to the exact transcriptional state of a patient’s tumour.

This naturally extends to artificial gene circuits for therapeutic implants. Imagine engineered cells implanted into the body that sense glucose, inflammation, oxygen, hormones, or disease-specific signals and respond by producing insulin, cytokines, antibodies, growth factors, or other therapeutic molecules. Today, this is still very hard because biological circuits are difficult to design, tune, and keep stable. But powerful DNA models could help design the regulatory DNA behind these systems: promoters, enhancers, repressors, feedback loops, insulators, and safety switches. In this scenario, DNA models would not only help us design single sequences; they would help us design living therapeutic devices that partially substitute for damaged cells, tissues, or organs.

Another possibility is trait borrowing across species. Evolution has already solved many biological problems: heat tolerance, cold resistance, drought survival, radiation resistance, unusual metabolism, immune evasion, regeneration, and extreme longevity. Future DNA models could help identify the genomic programs behind such traits and suggest how parts of them might be transferred, adapted, or reimplemented in other organisms. In agriculture, this could mean crops or livestock that are more resilient to climate stress. In biotechnology, it could mean microbes that borrow metabolic tricks from extremophiles. In medicine, it could mean discovering protective programs from species with unusual disease resistance.

There is also a future application in evolutionary forecasting. A DNA language model trained across huge evolutionary diversity might learn which mutations are plausible, which combinations are unstable, and which genetic routes are likely under selection. This could help anticipate viral escape, antibiotic resistance, cancer evolution, or the evolution of engineered organisms. Instead of only reconstructing the past, genomic models could help forecast the space of possible biological futures.

Pushed further, DNA models might become tools for simulating artificial evolution. In software, researchers are already exploring systems where agents write code, test it, mutate it, and improve it over many cycles. A similar idea could eventually be imagined for biology: populations of DNA sequences evolving inside computational environments, shaped by simulated selection pressures, constraints, and objectives. Such systems would not merely design one sequence at a time; they would explore evolutionary trajectories, helping us ask how new regulatory circuits emerge or how genomes adapt to new environments.

There is also the possibility of closed-loop biological design platforms. A model proposes DNA sequences, experiments test them, the results are fed back, and the model improves. Over time, this could create self-improving systems for regulatory element design, gene therapy optimization, microbial engineering, vaccine vector design, or cell therapy programming. The bottleneck would shift from “can we imagine the right sequence?” to “can we test enough sequences safely and learn from them efficiently?”

A very different frontier is DNA-based computation and data storage. DNA is already an attractive medium for information storage because it is dense, stable, and copyable. But practical DNA storage and molecular computing require solving hard design problems: how to encode information robustly, avoid problematic motifs, minimize synthesis and sequencing errors, retrieve specific files, and perform operations directly on molecules. Advanced DNA language models could help design better encoding schemes, error-correcting molecular formats, DNA circuits, and hybrid biological-computational systems.

The common theme across these applications is control. Today’s DNA models mostly help us read the genome. Tomorrow’s models may help us write it. The most transformative applications may come when models can connect sequence to function across scales: from nucleotides to regulatory elements, from regulatory elements to cell states, from cell states to tissues, and from tissues to disease, ecology, and evolution.

That is where DNA language models could become more than predictors. They could become design engines for biology.