Engineering Life From Scratch
Transcript
[SOFIA] So here's a fantasy that biologists have been chasing for decades: what if you could sit down at a keyboard and just write a genome? Type out the DNA you want, hit send, and get a living organism that runs your code.
[DANIEL] Which sounds absurd until you realize we're already partway there. We can synthesize DNA chemically — you can order oligos, short strands, off a website.
[SOFIA] Right, but the gap between "I can order a 200-base oligo" and "I have a functioning 178-kilobase megaplasmid inside a living cell" — that gap is enormous, and that's really the whole story today.
[DANIEL] Hm. Let me set the terms, because the audience might be coming from, say, structural biology and not know the vocabulary. When we say "writing a genome," we don't mean one uninterrupted print job. It's a stack of problems. You have to synthesize DNA, then assemble the small pieces into big pieces, then get those big pieces into a cell — and the cell has to tolerate them and copy them.
[SOFIA] And each of those steps has its own bottleneck. Synthesis is length-limited. Assembly gets messy when you're stitching lots of fragments. And delivery — getting DNA into a cell — is wildly organism-dependent. What works for E. coli falls flat in almost everything else.
[DANIEL] That last one matters more than people outside the field appreciate. For bacteria and archaea, getting DNA in is transformation — chemical or electroporation — or conjugation, where one cell physically passes DNA to another. And "the DNA replicates" requires an origin of replication the host recognizes. Miss any of that and your beautiful synthetic construct just... degrades.
[SOFIA] So the arc I want to trace is: how does a scrappy set of molecular tools slowly turn "writing DNA" into something you can actually do in weird organisms — not just E. coli. And the first turning point, for me, is 2022, in Deinococcus radiodurans.
[DANIEL] The radiation-resistant one.
[SOFIA] The absurdly radiation-resistant one. It shrugs off doses that would shred a human genome. Which makes it a fascinating chassis — but also a nightmare to engineer, because chemical transformation basically fails for large constructs.
[DANIEL] So what did they do?
[SOFIA] They used conjugation from E. coli. And this is the good stuff — they showed conjugation delivers both replicating and nonreplicating plasmids into Deinococcus where chemistry chokes. Then they did sequential gene deletions, swapping out four restriction-modification genes for resistance markers.
[DANIEL] The restriction-modification genes being the cell's immune system against foreign DNA — it chops up anything that doesn't carry the right methylation marks. Removing those is how you make a strain that'll actually accept your engineering.
[SOFIA] And then the move that made me sit up — they cloned the entire MP1 megaplasmid. 178 kilobases, about 62% G+C, notoriously hard to handle, captured with a nonreplicating plasmid carrying just one kilobase of homology to a target, pulled back into E. coli as a roughly 190-kb clone, and verified on a MinION.
[DANIEL] That's the direction people don't usually think about. Not writing DNA into the organism — reading a giant chunk out of it and stabilizing it in a lab workhorse. Both directions are "genome writing" infrastructure.
[SOFIA] Exactly. You can't rewrite what you can't capture and manipulate. So 2022 gives us delivery and capture at genome-fragment scale in a hard organism.
[DANIEL] But conjugation and homologous recombination, the way they did it there, are slow and low-throughput. Which is where 2023 pushes hard on efficiency. ORBIT.
[SOFIA] ORBIT is clever. So classic recombineering — lambda Red — lets you swap in DNA using short homology arms, but it's inefficient for big payloads. ORBIT splits the job. Lambda Red installs just a tiny landing pad, a Bxb1 attP site, carried on a targeting oligo. Then Bxb1 integrase — a serine integrase, an enzyme that does clean site-specific recombination — drops in a kilobase payload at that site.
[DANIEL] And the number they report is the eye-opener. A thousandfold over lambda Red for that payload delivery. Scaling to thirty-thousand-member libraries.
[SOFIA] A thousandfold! That's not a tweak, that's a different regime of what's feasible.
[DANIEL] It is — and I'd flag that this is E. coli. The thousandfold is in the workhorse, where everything already works well. The open question is portability. But as a proof that integrases can do the heavy lifting recombineering couldn't, it holds up beautifully.
[SOFIA] And that integrase theme keeps coming back. Which brings us to CAST, later in 2024 — CRISPR-associated transposases. Same underlying dream: put big DNA exactly where you want it, but here without a double-strand break.
[DANIEL] That "without a break" is the important part. Most targeted editing cuts the DNA and relies on the cell to repair it, which is lossy and can be lethal at scale. CASTs use an RNA guide for targeting, like CRISPR, but then a transposase inserts the payload — no cut.
[SOFIA] And Banta and colleagues turned that into genome-scale libraries in bacteria — knockouts, overexpression, protein fusions — all by programmable insertion. That's functional genomics at a throughput you couldn't touch with break-and-repair.
[DANIEL] So notice the through-line. Deinococcus: get DNA in and out at all. ORBIT: get it in efficiently at one site. CAST: get it in efficiently at many programmable sites, across a whole genome. Each one attacks the delivery-and-placement bottleneck from a different angle.
[SOFIA] And parallel to placement, you've got the assembly problem — building the constructs in the first place. That's the In- and Out-Cloning paper, also 2024. This one's for the cloning nerds and I say that lovingly.
[DANIEL] Golden Gate assembly, modular cloning — MoClo. The idea is you snap standardized DNA parts together in one pot using type IIS restriction enzymes that cut outside their recognition site, leaving defined overhangs.
[SOFIA] Right, and the perennial headache is scars — leftover junk sequence at the junctions. They used SapI, which leaves three-nucleotide overhangs, so you can make transcription units that are scar-free and codon-based. And Out-Cloning generates the acceptor plasmid you need for your specific organism, on demand, from modular parts.
[DANIEL] Which is the quiet enabler. If you're writing genomes for non-model organisms, you can't rely on off-the-shelf vectors. Making the destination plasmid part of the modular system — that's the kind of infrastructure that doesn't get headlines but unblocks everyone.
[SOFIA] And the last piece, for me, closes a loop back to reading. Coaux-Seq. You barcode roughly three-kilobase genomic fragments from eleven different bacteria, drop them into an E. coli expression vector, and select across twenty auxotrophic knockout backgrounds — strains missing a gene they need to grow.
[DANIEL] Complementation. If a fragment restores growth, it carries a gene doing that job. And the barcode plus cheap BarSeq readout tells you which fragment, at scale.
[SOFIA] They validated fifty-three protein functions experimentally — including the first sulfate importer in the TauE family, which everyone thought was an exporter family.
[DANIEL] That's the payoff I like. Writing genomes isn't only construction — it's learning what the parts actually do, so your future designs aren't guesses.
[SOFIA] So where's it heading? We've got efficient placement, scarless assembly, capture of huge fragments, and functional annotation — the pieces of a real write-read-rewrite cycle in organisms that aren't E. coli.
[DANIEL] The honest gap is integration and portability. Most of these shine brightest in the workhorse. The frontier is chaining them together in something like Deinococcus, end to end.
[SOFIA] Which is exactly the kind of ugly, thrilling engineering problem I'd love to cover once someone pulls it off. Daniel, take us out.
[DANIEL] Read the DNA, write the DNA, check your work. We'll be here when the field closes that loop.