CULTIVARIUM · RADIO
← On air
The Arc

Genomic Language Models Rewrite Biology

The Arc · with Sofia & Daniel · Recorded Sep 29, 2026
More episodes → Share on X Read the paper →
Transcript

[SOFIA] Okay, so here's a question that sounds almost silly until you sit with it: can you teach a machine to *read* DNA the way it reads English? Not just spell-check the letters — actually pick up grammar, syntax, meaning. That's the arc today. How we went from painstakingly hand-calling mutations to language models that treat genomes like text.

[DANIEL] And I want to flag up front — "language model" is doing a lot of work in that sentence. We should define it, because the whole story turns on whether that analogy actually holds.

[SOFIA] Right. So for anyone whose PhD is in, I don't know, membrane biophysics — a language model, in the machine learning sense, is a model you train by hiding parts of a sequence and asking it to predict what's missing. Mask a word, guess the word. Do that a few trillion times and the model builds up an internal sense of what tends to go where.

[DANIEL] The key move is it's self-supervised. Nobody hand-labels the data. You just feed it raw sequence, and the training signal comes from the sequence predicting itself. That matters for biology because labels are the bottleneck — we have oceans of sequence and a thimble of annotation.

[SOFIA] And that's the whole promise. We're drowning in genomes. Metagenomics especially — you sequence a scoop of soil, you get millions of genes, and most of them nobody's ever characterized. If a model could learn the "grammar" of DNA and protein from all that unlabeled data, you could start predicting function without doing the experiment first.

[DANIEL] Which is the dream. But let's earn it. Because before anyone was training on trillions of base pairs, the "machine reading DNA" job looked completely different. It was careful, hand-built variant calling. That's really where this starts.

[SOFIA] The 2015 mitochondrial cancer paper. And honestly I love starting here because it's the least AI thing on the list — and it makes the point about what "reading" DNA even means.

[DANIEL] So what they did: 527 tumors from TCGA, fourteen cancer types, and they called somatic mitochondrial mutations from both the DNA and the RNA of the same samples. Mitochondrial DNA is a nice testbed — tons of copies per cell, so you get enormous coverage. They were over five thousat-fold mean depth. That's how you get 616 high-confidence mutations you actually believe.

[SOFIA] And here's the part that gets me — they compared the DNA allele frequency to the RNA allele frequency, expecting them to just match. Correlation of 0.91. Beautiful. Except for fifteen outliers.

[DANIEL] And the outliers are the whole story. Twelve of the fifteen were transfer-RNA mutations. All of them sat in the stems of the tRNA — the base-paired parts of that cloverleaf fold — and they disrupted the predicted structure. And you saw unprocessed precursor RNA piling up, up to a hundred-fold.

[SOFIA] So the sequence wasn't the signal. The *shape* was. The cell reads the fold of the tRNA to know where to cut and process it, and if you break the stem, the whole neighborhood jams — the genes next door don't get matured either.

[DANIEL] They even worked out the clusters get processed three-prime to five-prime. Which is a mechanistic detail, but the deep point is: meaning in the genome isn't only in the letters. It's in the structure the letters fold into. Hold that thought, because every language model after this is basically trying to learn that lesson without being told.

[SOFIA] That's the through-line. Okay, jump to 2022 — DeepLoc 2.0. This is where the "language model" idea shows up in biology in earnest, on the protein side.

[DANIEL] Protein language models. Same recipe — mask amino acids, predict them, over millions of protein sequences. Models like ESM1b and ProtT5. DeepLoc took those learned embeddings and predicted where a protein goes in the cell — ten subcellular locations, multi-label, so a protein can go two places at once. And nine types of sorting signal.

[SOFIA] And this is the good stuff — they looked at the attention. The model's attention peaks landed right on the actual sorting signals. The little address labels proteins carry. Nobody told it "here's the signal peptide." It found them.

[DANIEL] Which is the protein echo of the 2015 mito result. The information was latent in the sequence, and a model trained only to predict masked residues learned to point at the functional part. I'll admit, when the attention co-locates with a known biological signal, that's the kind of thing that survives my skepticism. It's not proof of mechanism, but it's not nothing.

[SOFIA] So proteins fell first. And the obvious next question — why stop at protein? Genes live in context. They sit next to each other, they regulate each other. Which brings us to 2024, the OMG dataset and gLM2.

[DANIEL] This is the ambitious one. They built a corpus — 3.1 trillion base pairs, pulling metagenomes from JGI's IMG and EMBL's MGnify. And gLM2 is what they call mixed-modality. Instead of reading only amino acids or only nucleotides, it reads a whole locus — multiple genes in order, in both alphabets at once.

[SOFIA] Which is such a nice idea because a protein-only model literally cannot see the DNA between the genes. The promoters, the regulatory syntax, the spacing. gLM2 could, and they argue it picks up coevolutionary signals — like which protein interfaces evolve together — that you'd miss reading proteins in isolation.

[DANIEL] And they had to be clever to get there. Metagenomic data is wildly redundant — the same organism sequenced a thousand times. They deduplicated in embedding space rather than by sequence identity, so you're not just training on the same abundant bug over and over. That's a real methods contribution, not a footnote.

[SOFIA] And then 2025 kind of explodes in a few directions at once. WGRL goes whole-genome — learns embeddings from ordered conserved elements and beats the old protein-domain-presence-absence approach on 23 of 25 bacterial phenotypes.

[DANIEL] And I want to be precise there, because this one's listed as both supporting *and* in tension with the genomic-foundation-model idea. Beating domain presence/absence on 23 of 25 is a strong result. But the tension is real — bigger, general genomic models don't automatically win. Sometimes a focused representation of the right features beats the giant foundation model. That's a healthy fight for the field to be having.

[SOFIA] Then CodonTransformer — trained on a million gene pairs across 164 species — learns organism-specific codon usage. And the emergent bit: it starts avoiding negative cis-regulatory elements without being told to.

[DANIEL] "Emergent" is a word I always poke at. But operationally it means the behavior wasn't in the objective — they trained it to match codon distributions, and dodging bad regulatory motifs fell out for free. For someone like Sofia who wants to actually build genes for a non-model host, that's directly useful.

[SOFIA] Oh, hugely. That's a design tool. And the enzyme-annotation review ties the bow — it says we've gone from hand-crafted features to deep learning to, next, generative models that propose genuinely *new* functions instead of just relabeling known ones.

[DANIEL] Which is the whole trajectory in one line. 2015, we hand-called variants and inferred structure ourselves. 2025, we're asking models to read genomes in context and imagine new biology. The open question is still the 2015 question — do these models truly learn structure and function, or very sophisticated correlations?

[SOFIA] And that's the fight worth watching. Teaching machines to read DNA — we're past the alphabet. We're arguing about grammar. Daniel, thanks for keeping the controls honest.

[DANIEL] Always. Next segment after the break.