Machines Learn Genome Language
Transcript
[THEO] Okay, picture a language you can't read, written in just four letters, and it's billions of letters long with no spaces, no punctuation, and most of it looks like gibberish until suddenly it doesn't. That's a genome. And for a long time, we read it the way you'd read a foreign phone book — looking up one entry at a time.
[DR. MARA] Right. The classic approach was: find a gene, translate it to protein, match that protein against a database of things we already understand. BLAST, Pfam, domain annotation. It works beautifully — as long as the thing you're looking at resembles something we've seen before.
[THEO] And that's the catch for us at The Dish, because the organisms we care about — the weird environmental microbes, the non-model stuff — a huge fraction of their genes come back "hypothetical protein." Function unknown.
[DR. MARA] Dark genome. Depending on the lineage, you can have thirty, forty percent of genes with no confident annotation. The dictionary runs out. So the real question underneath this whole show is: can a machine learn to read DNA the way a large language model learns to read English — from the statistics of the sequence itself, without us hand-labeling every word?
[THEO] And that's a genuinely different idea. A language model doesn't memorize definitions. It learns that certain words show up near each other, in certain orders, and from that it builds an internal sense of grammar. The bet is that biology has a grammar too.
[DR. MARA] It does. Codon usage, regulatory motifs, the way genes cluster in operons, the way a protein folds. All of that constrains the sequence. The question was whether a model could pick it up unsupervised. And I'd argue the roots of this story go back before anyone said "foundation model" — back to people noticing that structure, not just letters, carries the signal.
[THEO] This is the 2015 mitochondrial paper, right? That one surprised me as a starting point.
[DR. MARA] It's a lovely origin point precisely because there's no machine learning in it. They took 527 tumors from TCGA, called somatic mitochondrial mutations from DNA and RNA at over five thousand-fold coverage, and asked: do the DNA and RNA allele frequencies agree? Mostly, yes — correlation of 0.91.
[THEO] But the outliers are the whole story.
[DR. MARA] Fifteen outliers. Twelve of them were mitochondrial tRNA mutations. And every one sat in the stem of the tRNA's folded structure, broke the predicted fold, and the cell couldn't process it — the precursor piled up, unspliced, up to a hundredfold.
[THEO] So the cell's machinery wasn't reading the sequence of the tRNA. It was reading the shape. Mess up the fold and the whole processing line jams — even the neighboring genes don't get matured.
[DR. MARA] They even saw the clusters get processed in a preferred three-prime to five-prime direction. The lesson that carries forward: the meaning isn't in the letters alone. It's in what the letters make the molecule do. Any model that only reads sequence as text is going to miss that unless the structure leaves a statistical fingerprint.
[THEO] Which is the thread that runs all the way through. Okay — jump to 2022, DeepLoc 2.0. First time we really see the language-model machinery show up.
[DR. MARA] This is proteins, not DNA yet. They take a protein language model — ESM1b, ProtT5, these are models trained on hundreds of millions of protein sequences — and they ask: where in the cell does this protein go? Ten possible locations, and a protein can have more than one, so it's multi-label.
[THEO] And the part I love — the attention. The model learns to "look at" certain stretches of the sequence, and when they checked where it was looking, the attention peaks landed right on the sorting signals. The little address labels that tell a protein where to go.
[DR. MARA] Without being told those signals existed. That's the proof of concept. A model trained just to understand protein sequences had internalized the targeting grammar on its own. You could open the hood and see it pointing at the right residues.
[THEO] So that's the "it actually learned real biology" moment.
[DR. MARA] For proteins. The next leap is going back to DNA — and not throwing away everything proteins taught us. That's the 2024 OMG dataset and the gLM2 model.
[THEO] This is the big one for the non-model world. Three-point-one trillion base pairs of metagenomics — JGI and MGnify stitched together — environmental DNA from everywhere.
[DR. MARA] And the clever move is "mixed-modality." Earlier models were either nucleotide or amino acid. gLM2 reads a whole genomic locus as it actually sits on the chromosome — multiple genes in a row — and lets the model see amino acids and nucleotides together.
[THEO] Why does that matter? Because the protein-coding part has one grammar, but the stuff between genes — promoters, operators, spacing — that's nucleotide grammar. If you only read protein, you're blind to the regulation.
[DR. MARA] Exactly. And because it reads neighboring genes together, it picks up coevolution — when two proteins touch, their interfaces change in a correlated way across evolution. A protein-only model looking at one sequence at a time can't see that. gLM2 can.
[THEO] One quick aside on scale — they had to deduplicate three trillion base pairs. They did it in embedding space, meaning they let the model's own representation decide what counted as redundant. Which is very meta.
[DR. MARA] Necessary, too, or the model just memorizes the overrepresented genomes. Now — 2025 is where it gets contentious, and I like that it gets contentious. WGRL.
[THEO] Whole-genome representation learning. And this one's got a little feud baked into it.
[DR. MARA] It does. They train a long-context language model on ordered conserved elements across a whole genome, and the embeddings beat the old protein-domain presence/absence method on 23 of 25 bacterial phenotypes.
[THEO] So "does this genome have these protein domains, yes or no" — the honest, interpretable workhorse — loses to the learned whole-genome representation on almost everything.
[DR. MARA] And yet the brief flags it as in tension with the genomic-foundation-model idea at the same time it builds on it. Which is the real state of the field: these models work, but how much is genuine genomic grammar versus a very good lookup of conserved elements is still argued. I'd want to see the controls before I call it understanding.
[THEO] Then CodonTransformer, same year, is the generative turn.
[DR. MARA] A million gene pairs across 164 species. It learns organism-specific codon grammar — how a given species prefers to spell the same amino acid — and generates sequences matching the natural distribution. The striking part: it emergently avoids negative cis-regulatory elements it was never explicitly told to avoid.
[THEO] Which is DeepLoc's attention trick all over again — the model picks up a rule nobody wrote down. And that's enormous for expressing a gene in a weird host.
[DR. MARA] It's the design side of reading. And the 2025 enzyme-annotation review names the destination plainly: we're moving from using these models to re-label things we already know toward generating genuinely novel functions.
[THEO] So the arc: 2015 says the signal is in the structure, not the letters. 2022 shows a model can find the signal itself. 2024 reads DNA and protein together across the environmental unknown. 2025 argues about how much it really understands — and starts writing new sequences.
[DR. MARA] From reading one word at a time to drafting new sentences. Still with a healthy fight about whether the machine knows what it's saying.
[THEO] Which is the best kind of open question. More on that after the break.