CULTIVARIUM · RADIO
← On air
The Arc

Genomic Language Models Decoding Biology

The Arc · with Sofia & Daniel · Recorded Oct 4, 2026
More episodes → Share on X Read the paper →
Transcript

[SOFIA] Okay, so here's a question that sounds almost too simple: can a machine learn to read DNA the way a model reads language? Not just spell-check it — actually understand the grammar, the dialect, what a sentence means in context.

[DANIEL] Hm. And "read" is doing a lot of work in that sentence.

[SOFIA] It is! That's exactly what I want to dig into today. Because over the last decade the field has gone from very careful, hand-built biology to these enormous models that claim to learn the language of genomes on their own. And I think the arc tells you something about what we actually mean by "understanding" a genome.

[DANIEL] Let me set the table for anyone who doesn't live in this corner. When we say "language model" here, we mean the same machinery behind the text models everyone's used — you train on a massive pile of sequence, you mask out pieces, and the model learns to predict what's missing. Do that at scale and it builds internal representations — embeddings — that capture statistical regularities. For text, that's syntax and meaning. For biology, the hope is it captures structure, function, regulation.

[SOFIA] And the unit matters. You can feed it protein sequence — strings of amino acids. Or raw nucleotides, the A-C-G-T. Those are different alphabets with different grammars.

[DANIEL] Right, and that distinction becomes the whole story later. But I want to start somewhere that isn't a language model at all — because the roots here are old-fashioned, careful molecular biology.

[SOFIA] The 2015 mitochondrial paper.

[DANIEL] Yes. And on the surface it has nothing to do with AI. The group took 527 tumors from TCGA, fourteen cancer types, and called somatic mutations in the mitochondrial genome — but they did it from both DNA and RNA. Paired whole-genome sequencing and polyA RNA-seq. Mitochondria give you absurd coverage, over five thousand-fold mean, so you can call low-frequency variants with confidence. 616 high-confidence mutations.

[SOFIA] And the DNA and RNA frequencies tracked beautifully — correlation of 0.91.

[DANIEL] They did. Which is almost boring, except for the outliers. Fifteen mutations where DNA and RNA disagreed, and twelve of those — eighty percent — were in mitochondrial tRNAs. Every one of them sat in a stem of the cloverleaf fold. They broke the predicted secondary structure, and the unprocessed precursor piled up — up to a hundred-fold.

[SOFIA] This is the part I love. The cell isn't reading the tRNA's sequence to decide how to process it. It's reading the shape.

[DANIEL] That's the claim, and the data support it cleanly. Mutations that changed sequence but preserved the fold were fine. Mutations that wrecked the fold blocked maturation — and not just of that tRNA, of its neighbors in the cluster. They even saw directional processing, three-prime to five-prime.

[SOFIA] So why is this the root of a show about AI reading DNA?

[DANIEL] Because it's the thesis in miniature. The information that matters isn't in the letters. It's in the structure the letters fold into, and in the context — what sits next to what. A model that only reads sequence token by token would miss exactly the signal that mattered here.

[SOFIA] That's a great way to frame it. The 2015 people found that signal by hand — predicting folds, lining up precursors. The whole arc after this is machines trying to learn that lesson on their own.

[DANIEL] And the first real step toward that is DeepLoc 2.0, in 2022.

[SOFIA] Okay, this one's protein. The problem is subcellular localization — where in the cell does a protein go? Nucleus, mitochondrion, secreted, membrane. And proteins carry little address labels, sorting signals, short stretches that say "ship me here."

[DANIEL] And DeepLoc doesn't engineer features for those signals. It takes embeddings from protein language models — ESM1b, ProtT5 — models already trained on huge protein databases, and it pools them with an attention mechanism. Predicts ten locations, multi-label, because a protein can live in two places. Plus nine types of sorting signal.

[SOFIA] And here's the good stuff — when they looked at where the attention landed, the peaks sat right on top of the real sorting signals.

[DANIEL] That's the result that should make you sit up. Nobody told it where the signals were. It learned that those residues carried the information. That's the 2015 lesson arriving in machine form — the model found the functional element in context.

[SOFIA] So from there the ambition explodes. If a protein model can do that, why stop at proteins?

[DANIEL] Which brings us to 2024 and the OMG dataset — gLM2. This is the turning point for me. They assembled 3.1 trillion base pairs of open metagenomic data, JGI and MGnify combined, and trained what they call a mixed-modality model.

[SOFIA] Mixed-modality — meaning it reads amino acids and nucleotides in the same model, laid out the way they actually sit in a genome. Genes in order, with the stuff between them.

[DANIEL] That's the key move. A protein-only model sees each protein in isolation. gLM2 sees a whole locus — this gene, then the intergenic region, then the next gene. And they report it learns coevolutionary signals across protein interfaces and regulatory syntax that protein-only models structurally cannot see.

[SOFIA] Because the context is the point. Proteins that touch each other evolve together, and if you only ever see them apart you lose that correlation.

[DANIEL] Now, I'll flag — "learns coevolutionary interface signals" is the kind of claim I want to see probed hard. But the architecture is the honest advance: context across the whole locus, not one gene at a time.

[SOFIA] And 2025 is where the field starts arguing with itself, which I find delightful.

[DANIEL] It does. WGRL — whole-genome representation learning. Self-supervised language modeling, but over ordered conserved elements across a whole bacterial genome, to produce one embedding per genome. And they test it on 25 bacterial phenotypes against a classic baseline — protein domain presence/absence, basically "which Pfam domains does this genome have."

[SOFIA] The old reliable. Count the parts.

[DANIEL] And WGRL beats it on 23 of 25. Which is a real result. But notice it's listed as in tension with genomic foundation models — the generic "train a giant model on raw sequence" recipe. WGRL's bet is that order and arrangement carry phenotype, not just raw nucleotide prediction at scale.

[SOFIA] So that's the live argument: do you just make the model bigger and feed it everything, or do you structure what you feed it — genes in order, conserved elements, mixed modality?

[DANIEL] Exactly the fault line. And CodonTransformer, also 2025, is the generative flavor. They trained on a million gene pairs across 164 species, with this STREAM masking scheme, and the model learns organism-specific codon grammar — how a given species prefers to spell the same protein.

[SOFIA] And the emergent bit — it avoids negative cis-regulatory elements without being told to. It just picks codon choices that don't accidentally spell a problematic signal.

[DANIEL] Which for your world is directly useful.

[SOFIA] Hugely. If I'm trying to express a gene in a non-model organism, I need the codons to match that host's dialect. A model that writes in the native accent and dodges bad regulatory motifs — that's a design tool I'd use tomorrow.

[DANIEL] And the Kim review on enzyme annotation closes the arc honestly. It traces the same trajectory — hand-crafted features, then deep learning that extracts them automatically — and argues the frontier is generative AI finding genuinely new enzyme functions, not just relabeling known ones.

[SOFIA] Which is the whole arc in one line. We started with people reading structure by hand in mitochondrial tRNAs. We're ending with models that read context, read dialect, and now want to write new biology.

[DANIEL] And the open question is whether "reads context well" means "understands" — or just predicts. I'd want the controls before I'd say understands.

[SOFIA] Fair. But the direction's clear, and it's the context that got us here. Daniel, great map — let's take a breath, and when we're back, one organism that's about to feel all of this.