AI Automates The Experimental Loop
Transcript
[SOFIA] Okay, so here's a question I think about a lot. When you picture AI changing biology, what do you see? Because I bet most people picture a chatbot writing up your methods section, or maybe AlphaFold spitting out a protein structure. Clean. Digital. On a screen.
[DANIEL] Which is the easy part, honestly. Prediction is cheap. The hard part has always been the wet lab.
[SOFIA] Right! And that's the thread I want to pull today — AI at the lab bench. Not AI as a smarter autocomplete, but AI that actually touches the experiment. That decides what to grow, what to build, what to test next. And it turns out that story has roots going back just a couple of years, and it's moving fast.
[DANIEL] So let's set the problem up, because I think it's underappreciated. Biology runs on a loop. In synthetic biology people call it DBTL — Design, Build, Test, Learn. You design a construct, you build it, you test it, you learn something, and you go around again.
[SOFIA] And that loop is slow and expensive. Each turn can be weeks. The organism doesn't care about your deadline.
[DANIEL] And every turn you're making decisions — which media, which variant, which gene — mostly on intuition and whatever the last experiment told you. The space of things you could try is astronomically larger than the number you can actually run.
[SOFIA] So the dream is: can a machine help you search that space smarter? And there are really two flavors of this. One is AI that physically runs the loop — robots, imaging, automation. The other is AI that reasons about biological sequence — what protein should I even make. Both are "AI at the bench," and they start converging. Let me define a couple terms for anyone outside the field.
[DANIEL] Please, because we're going to say "foundation model" a lot.
[SOFIA] A foundation model is a big neural network trained on a huge pile of data in a general way — like the language models everyone knows, but here the "language" is DNA or protein sequence. You train it to predict the next token across millions of sequences, and it picks up the grammar of biology without being told the rules. Then you can fine-tune it for a specific job.
[DANIEL] And the counterpart to all this fancy modeling is directed evolution — the old reliable. You take an enzyme, you mutate it, you screen for the ones that work a little better, you repeat. It won a Nobel Prize. It works. But it's local search — you can only climb the hill you're standing on.
[SOFIA] That tension — local hill-climbing versus a model that can jump across the landscape — that's going to come back. Okay. Roots. Where does this arc start?
[DANIEL] The humblest place possible. 2023, a paper called SPIRO — the Smart Plate Imaging Robot. It's a Raspberry Pi and some 3D-printed parts that photographs Petri plates on a schedule.
[SOFIA] And I love that it starts here, because it's not glamorous and that's the point. SPIRO lets a plant biologist with zero engineering background automate germination and root-growth phenotyping at scale. And here's the clever bit — it can image in the dark, because it controls its own lighting. Root biology in true darkness was basically impractical before.
[DANIEL] It's the unsexy foundation. Before AI decides anything, something has to generate clean, consistent data. SPIRO is the data-generation layer. No model survives contact with messy, hand-collected measurements.
[SOFIA] And the same year — this is the turning point for me — BacterAI. Same idea of automation, but now there's a brain making decisions. The problem they tackled: mapping what a microbe needs to grow. Which of dozens of nutrients can you remove and the bug still lives?
[DANIEL] Which is a brutal combinatorial problem. Every ingredient on or off — you can't test all combinations.
[SOFIA] So they reframed it as a reinforcement learning problem. A Markov decision process — the agent gets rewarded for removing ingredients and still seeing growth. So it pushes its own experiments right up to the edge, the boundary between growth and no-growth, which is exactly where you learn the most.
[DANIEL] And that's the elegant part. It's not screening blindly. It's steering toward the informative experiments, and it mapped auxotrophies — the nutrients the organism can't make itself — in days instead of a career.
[SOFIA] That's the first real "AI at the bench." A closed loop: the model proposes, the robot runs it, the result trains the model, overnight. Okay, this is the good stuff — because once you've shown a machine can drive the loop, the question becomes what else can it drive.
[DANIEL] And this is where the sequence-modeling branch comes roaring in, 2025. GenomeOcean — a four-billion-parameter genome foundation model, trained not on tidy reference genomes but on metagenome co-assemblies. Raw environmental DNA.
[SOFIA] Which matters because that's where the weird biology lives — the stuff from organisms nobody's cultured. They used a smart tokenization scheme, byte-pair encoding, and the headline is it generates protein-coding sequence about 150 times faster than a comparable model, Evo-7B.
[DANIEL] And zero-shot discovery of biosynthetic gene clusters — BGCs, the stretches of DNA that encode antibiotics and other natural products. Fine-tuned into a version they call bgcFM, it finds novel ones it was never explicitly trained to find.
[SOFIA] And on the protein side you get Germinal, which I think is gorgeous. Designing antibodies — specifically the CDRs, the little loops that actually grab the target. It co-optimizes two things at once: AlphaFold-Multimer's structural confidence, is this fold real, and an antibody language model saying, does this look like a natural antibody.
[DANIEL] Nanomolar binders against four different protein targets, in something like 43 to 101 designs. That's a believable number — that's the part that makes me sit up. Small enough you can actually test them all.
[SOFIA] And then the two 2025 vision papers tie the bow. One argues for unified generative models that encode sequence, structure, and function together — basically, genetically encode any chemistry you want.
[DANIEL] And here's the clash I want to flag. That enzyme-discovery vision is explicitly in tension with machine-learning-guided directed evolution. The directed-evolution camp says: stay close to things that work, make small reliable steps. The generative camp says: the whole point is to escape local search and jump somewhere new. Those are genuinely different bets about where good enzymes come from.
[SOFIA] And the last paper reorders the whole loop. LDBT — Learn, Design, Build, Test. Instead of designing first and learning last, you let a model make a zero-shot prediction before you design anything, and you pair it with cell-free synthesis — expression without living cells — so you can test at massive scale.
[DANIEL] Toward single-round campaigns. One trip around the loop instead of ten. Which is the whole arc in one sentence — SPIRO automated the looking, BacterAI automated the deciding, and LDBT wants to collapse the loop so you barely go around it.
[SOFIA] From a Raspberry Pi watching seeds sprout to a model that proposes the experiment and a cell-free system that runs a million of them. Same instinct the whole way: let the machine search the space you can't.
[DANIEL] And I'd just keep my skeptic's hat on. The visions are visions — the proof is still a binder you can measure or a strain that grows. GenomeOcean and Germinal put real numbers on the table. The single-round dream hasn't yet.
[SOFIA] Fair. That's the next show, honestly. For now — that's the arc. Daniel, thanks for keeping me honest.
[DANIEL] Always.