AI Steals The Lab Bench Decision
Transcript
[SOFIA] Okay, so I want to start with a picture. Picture a lab bench — the one you did your PhD at. There's a stack of plates, a pipette, a notebook with your terrible handwriting. Now imagine the thing deciding what experiment to run next isn't you. It's a model. That's the arc we're tracing today — how AI stopped being the thing that analyzes your data after the fact, and started being the thing that picks the next move.
[DANIEL] Which is a bigger shift than it sounds. For most of the last decade, "AI in biology" meant something fairly narrow — you'd run your experiment, generate a pile of sequencing reads or images, and a model would classify or cluster it. The intelligence lived at the end of the pipeline.
[SOFIA] The read-out, not the decision.
[DANIEL] Right. And what's happening across these papers is the model creeping earlier and earlier — into the design, into which experiment even gets done. The whole loop.
[SOFIA] And the loop has a name people in synthetic biology will know — Design, Build, Test, Learn. DBTL. You design a construct, you build it, you test it, you learn from the result, and then you go around again. The problem is each lap is slow and expensive. You might do three, four cycles in a year if you're lucky.
[DANIEL] And every cycle, you're exploring a tiny neighborhood of a gigantic space. If you're engineering an enzyme, the number of possible sequences is astronomically larger than anything you can physically test. So the real question underneath all of this — how do you search a space you can never fully sample?
[SOFIA] That's the through-line. Hold onto that. Let me define a couple of terms for anyone coming from a different field. Directed evolution — you take a protein, make a bunch of random variants, screen for the ones that work a little better, repeat. It won a Nobel. It works. But it's local search. You're climbing the hill right next to where you're standing.
[DANIEL] You can't jump to a different mountain. And a foundation model — the other term worth defining — is a big model trained on enormous amounts of data, in this case sequences, so that it learns the general statistics of what real biology looks like. Then you can point it at a specific task without retraining from scratch. Zero-shot, when it does the task cold, with no task-specific examples.
[SOFIA] Okay. So that's the setup — slow loops, impossibly big spaces, and this new idea that a model that's "seen" enough biology might shortcut the search. Where do we start the story?
[DANIEL] I'd start somewhere humble, honestly. 2023, the SPIRO paper. It's a 3D-printed Raspberry Pi plate imager. No neural net doing anything fancy.
[SOFIA] And I love that you want to start there, because it makes the point. Before AI can pick your experiments, something has to generate data at a scale and consistency a human can't. SPIRO is a little robot that photographs Petri plates on a schedule — germination, root growth — and crucially it can image in the dark, which you physically can't do by walking in with your eyeballs because you'd turn the light on.
[DANIEL] And it's built for biologists with no engineering background. That's the quiet turning point. The bottleneck wasn't ideas, it was that automation used to require an engineer in the room. SPIRO says — here's open hardware, print it, run it. The bench starts feeding the model.
[SOFIA] Then same year, BacterAI. And this one, okay, this is the good stuff. They're mapping what a bacterium needs to grow — which nutrients it can't make itself, its auxotrophies. Normally brutal combinatorics. Dozens of ingredients, you can't test every combination.
[DANIEL] And instead of brute force, they frame it as a reinforcement learning problem — a Markov decision process. The agent gets rewarded for removing ingredients while keeping the cells alive.
[SOFIA] So it's actively pushing right up against the edge of what the bug can survive on.
[DANIEL] The growth front. And that's the elegant part — the experiments it chooses to run are the maximally informative ones, right on the boundary. It learns the auxotrophies in days. This is the model picking experiments, in the physical world, closing the loop autonomously. That's the real debut of the idea for me.
[SOFIA] So now watch the idea split into two branches. One branch is the generative-design dream. The 2025 enzyme-discovery vision paper lays it out — a single unified model that encodes sequence, structure, and function together, so you could, in principle, genetically encode any chemistry you want. Not climb the local hill. Just... ask for the mountain.
[DANIEL] And I want to flag — that one's framed as a vision. It's the aspiration, and the brief itself notes it sits in tension with machine-learning-guided directed evolution.
[SOFIA] Say more about that tension, because I think it's the heart of the argument.
[DANIEL] ML-directed evolution still keeps the evolutionary loop — the model just guides which variants you try next. It's smarter local search. The generative vision wants to skip the search. And those are genuinely competing philosophies about where the intelligence should sit. Do you trust the model to propose the answer, or do you keep evolution in the loop as a reality check?
[SOFIA] And the newer papers are basically evidence in that argument. GenomeOcean — a four-billion-parameter model trained on metagenome co-assembly. Trained on DNA straight out of environmental samples, not just tidy reference genomes.
[DANIEL] Which matters because that's where the undiscovered biology actually lives. And two concrete things: it generates protein-coding sequence about 150 times faster than the comparable Evo-7B model, and after fine-tuning — they call it bgcFM — it finds biosynthetic gene clusters zero-shot.
[SOFIA] BGCs — stretches of genome that encode the machinery for making natural products. Antibiotics, lots of drugs come from these. Normally you hunt them with tools like antiSMASH that pattern-match against known clusters.
[DANIEL] And "zero-shot discovery" means it's flagging novel ones without being shown examples of that class. Now — the brief lists it as in tension with genomic foundation models as a concept, so even here the field isn't settled on whether these models generalize or just memorize.
[SOFIA] Then Germinal, which I think is the most concrete win. Antibody design. They co-optimize two models — AlphaFold-Multimer for whether the structure's confident, and an antibody language model, IgLM, for whether the sequence looks like a real antibody.
[DANIEL] And the numbers are what make me sit up. Nanomolar binders against four different protein targets, in 43 to 101 designs. Success rates four to twenty-two percent. That's a real, falsifiable, measured result — not a vision.
[SOFIA] And it answers your "where should the intelligence sit" question in a specific way — jointly. Structure confidence and naturalness, pulling against each other, keeps it honest.
[DANIEL] It's two constraints checking one another, which is exactly the kind of thing I want to see instead of one model's unverified confidence.
[SOFIA] And it all converges on the last paper — the LDBT paradigm. They literally reorder the letters. Learn, Design, Build, Test. Put the model's zero-shot prediction first, before you design anything, and pair it with cell-free synthesis so the Build and Test happen at massive scale without even growing cells.
[DANIEL] The ambition being to collapse those four slow laps a year into — ideally — a single round.
[SOFIA] So that's the arc. From a Raspberry Pi taking pictures in the dark, to a reinforcement agent at the growth front, to models that propose the molecule before you touch a pipette.
[DANIEL] With the open question still open — predict-and-propose versus evolve-and-guide. Germinal says keep two models honest. GenomeOcean says the tension's unresolved. I'd rather the argument stay live than get declared won.
[SOFIA] Honest and unfinished. That's a good place to leave it. More after the break.