Engineering Genomes From Scratch
Transcript
[SOFIA] Okay, so I want to start with a fantasy that's been floating around biology for decades. Imagine you don't just edit a genome — you write one. From scratch. You sit down at a keyboard, you type out the sequence you want, and something hands you back a living cell that runs on it.
[DANIEL] Hm. That's the dream people invoke when they say "synthetic genomics." And I'd push back a little right at the top — writing a whole genome from nothing is still extraordinarily hard. What's actually happening is a stack of enabling technologies, and each one solves a piece.
[SOFIA] Right, and that's the story I want to tell today — not one heroic paper, but the tool chain. Because the reason this matters isn't just "cool, we made a cell." It's that if you can write DNA reliably, you can engineer organisms that fix nitrogen, that eat plastic, that survive conditions no lab strain would tolerate. Building instead of tinkering.
[DANIEL] So let's define the problem for someone coming in from, say, immunology. Three things have to work. One — you need to physically get large pieces of DNA into a cell, which is nastier than it sounds the bigger the piece gets. Two — you need to assemble those pieces precisely, no leftover junk sequence at the joints. And three — you need to put the DNA at a specific place in the genome, cleanly.
[SOFIA] And the vocabulary trap I always flag — for bacteria, getting DNA in is called transformation, or conjugation, or electroporation. Not transfection. Transfection is a eukaryote word. Bacteria get transformed.
[DANIEL] Which brings us to the first paper in the arc, 2022, and it's a delivery paper. The organism is Deinococcus radiodurans — this is the bacterium famous for surviving radiation doses that would shred anything else. Charismatic, weird, and notoriously hard to work with.
[SOFIA] And that's exactly why it's a good test case! If your tools only work in nice, polite E. coli, you haven't really built tools for biology — you've built tools for one lab rat. Deinococcus has a high-GC genome, it's got these restriction-modification systems that chew up incoming foreign DNA —
[DANIEL] Those are the bacterial immune systems — restriction enzymes that cut DNA that isn't methylated the way the host expects. A defense against phage, but also a wall against anything you're trying to introduce.
[SOFIA] So chemical transformation just fails for big constructs. And what this group showed is that conjugation — literally mating the cells, E. coli passing DNA to Deinococcus through direct contact — gets both replicating and non-replicating plasmids across when the chemical route dies.
[DANIEL] And here's the part that made me sit up. They knocked out four of those restriction-modification genes, sequentially, by replacing each with a resistance marker. So they're using the delivery method to disarm the very defenses that block delivery. Bootstrapping.
[SOFIA] Okay, this is the good stuff — then they run it in reverse. They used a non-replicating plasmid carrying about a kilobase of homology, and they captured the entire 178-kilobase megaplasmid, this MP1, cloned it back into E. coli as a roughly 190-kb construct. Verified on a MinION.
[DANIEL] Which is the write-genomes-from-scratch dream in miniature, honestly. You're not editing a base here or there — you're moving a whole replicon between organisms. My skeptic's note: it's one megaplasmid, one demonstration. But the logic is sound and the sequencing backs it.
[SOFIA] So that's delivery and capture. The next problem is precision — putting a payload exactly where you want it. And 2023 gives us ORBIT for E. coli.
[DANIEL] So the old workhorse here is lambda Red recombineering — you use phage recombination proteins to swap short pieces of DNA into the chromosome. It works, but efficiency falls off a cliff as the insert gets bigger.
[SOFIA] ORBIT is such a clean trick. You use lambda Red to install just a tiny landing pad — a Bxb1 attP site, carried in on the targeting oligo. And then a serine integrase, Bxb1, does the heavy lifting of dropping in a kilobase payload at that site.
[DANIEL] And the number is the headline — a thousand-fold over lambda Red for the big payload, scaling to thirty-thousand-member libraries. That's the turning point in my mind. You've decoupled "mark the spot" from "deliver the cargo," and you let each enzyme do what it's good at.
[SOFIA] Which is so an engineer's instinct — modularity. Same year, and related in spirit, there's the ADDomer vaccine paper. Now this one's a bit of a side branch —
[DANIEL] It is. It's a eukaryotic-flavored, protein-nanoparticle story. Sixty copies of an adenovirus penton-base protomer self-assemble into a dodecahedron, thermostable, melting temp around fifty-five, and you can genetically insert a 33-residue epitope into an accessible loop.
[SOFIA] And what I love is the affinity math. They raised a nanobody by ribosome display that binds Omicron RBD at only 2.4 micromolar as a single monomer — weak! But display it thirty-six times on the same scaffold and you're at picomolar, a 42-picomolar cell-attachment EC50.
[DANIEL] Avidity. Many weak grips become one strong one. And the reason it belongs in this arc — it's writing sequence to program self-assembly. You're designing at the DNA level and getting structured matter out. Different domain, same philosophy.
[SOFIA] Then 2024 gets really deep into the assembly problem. In- and Out-Cloning, the SapI paper. So MoClo — modular cloning — is the Golden Gate world where you snap DNA parts together using type-IIS enzymes that cut outside their recognition site and leave defined overhangs.
[DANIEL] The nagging problem with Golden Gate has always been scars — little leftover sequences at the junctions. SapI leaves a three-nucleotide overhang, and because a codon is three nucleotides, you can make the joints land on codon boundaries and get scarless transcription units.
[SOFIA] And Out-Cloning is the bit I'd underline — it builds organism-specific acceptor plasmids on demand, for any organism, from modular parts. Which loops us straight back to Deinococcus, right? Non-model organisms all need their own vectors.
[DANIEL] Then CAST — RNA-guided targeted transposition. CRISPR-associated transposases. This is the natural successor to ORBIT's landing-pad idea, but now the targeting is programmable by a guide RNA, and crucially — no double-strand breaks.
[SOFIA] Which matters because double-strand breaks are toxic and unpredictable in a lot of bacteria. CASTs let you drop large payloads at a chosen locus and build genome-scale libraries — knockouts, overexpression, protein fusions — all without cutting.
[DANIEL] And the last one, Coaux-Seq, closes the loop on function. Because once you can write and place DNA, you have to ask — what does it do? They barcode roughly 3-kb genomic fragments from eleven diverse bacteria, express them in E. coli, and select across twenty auxotroph knockout backgrounds. Cheap BarSeq readout.
[SOFIA] And they validated 53 protein functions experimentally — including a TauE-family protein that turned out to be a sulfate importer, when everyone had it filed as an exporter family.
[DANIEL] That's the falsifiable payoff I like. Gain-of-function complementation — you don't just annotate by homology, you make the cell prove the gene does the job.
[SOFIA] So the arc — deliver it, place it, assemble it scarless, target it without breaks, and then test what it does at scale. That's the write-genomes toolkit assembling itself in real time.
[DANIEL] And it's heading toward non-model organisms specifically. The demonstrations keep leaving E. coli behind.
[SOFIA] Which is where the interesting biology lives. We'll pick up the environmental side after the break — stay with us.