CULTIVARIUM · RADIO
← On air
Today in AI

Today in AI — Aug 28

Today in AI · with Theo & Dr. Mara · Recorded Aug 28, 2026
More episodes → Share on X Read the paper →
Transcript

[THEO] Alright, picture this: you've got these incredibly powerful AI systems, right? They're getting smarter, faster, and able to do more complex things. But how do we *really* know if they're safe? How do we even know what they're truly capable of without bias creeping in?

[DR. MARA] That's precisely the challenge, Theo. Evaluating advanced AI, especially frontier models, isn't straightforward. Traditional benchmarks often fall short because they can be gamed, or they don't capture the nuanced capabilities or failure modes of these systems in real-world scenarios. It's like trying to assess a supercomputer with a calculator test.

[THEO] And that's why this announcement from Google DeepMind caught my eye. They're piloting what they're calling the "world's first double-blind evaluations" for frontier AI models. That's a huge methodological leap for evaluating these systems, isn't it?

[DR. MARA] It is. The "double-blind" aspect is crucial here. In a typical double-blind study, neither the participants nor the researchers know who is receiving the experimental treatment or the placebo. Applied to AI, it means that the human evaluators assessing the AI's output don't know which model generated the response – or even if it was AI-generated at all. And, crucially, the developers of the AI don't know which specific tests or prompts their model is being subjected to.

[THEO] So, it's like a scientific experiment for the AI itself, removing as many human biases as possible from the evaluation process. That's a big deal when you think about how quickly these models are evolving.

[DR. MARA] Exactly. It aims to provide a more objective assessment of capabilities, safety, and potential risks, without the influence of preconceived notions about a particular model or developer. This kind of rigor is standard in many scientific disciplines, but it's new for frontier AI evaluation, where the systems are so complex and the stakes are so high.

[THEO] Meanwhile, over at Anthropic, they're not just evaluating; they're pushing their Claude models into active roles in the lab. They've launched a research system where Claude can actually manipulate laboratory tools. That's a pretty tangible step into the physical world for an AI.

[DR. MARA] Indeed. This moves Claude from being purely a conversational agent or a research assistant to an active participant in experimentation. The idea is that an AI, following a "Model Hardware Standard," could directly control scientific and engineering devices to conduct experiments. This could significantly accelerate discovery by automating parts of the experimental cycle that are currently labor-intensive.

[THEO] Like having a really smart grad student who never sleeps.

[DR. MARA] A very precise one, yes. It also opens up possibilities for what they're calling "Claude for science," expanding support for researchers. This is a clear signal that these labs see AI as not just *about* science, but *doing* science.

[THEO] And speaking of increasing capabilities, the chatter online, and from OpenAI, is about a new pretraining run they’ve reportedly completed, codenamed "Bel," with a staggering number of parameters – over 10 trillion. If true, that's a significant jump.

[DR. MARA] The parameter count, while not the sole metric, generally correlates with increased model complexity and potential capabilities. A successful pretraining run of that magnitude suggests a substantial investment in scaling up their foundational models, potentially leading to more sophisticated reasoning and understanding. The online discussion also includes OpenAI noting they're making their custom inference chips, like "Jalapeño," faster and more energy efficient, which ties into the scaling of these massive models. More intelligence per watt, they're saying, and faster responses.

[THEO] So, we’ve got new methods for rigorous evaluation, AI taking a more active role in the lab, and models getting significantly larger and more efficient. It feels like the toolkit for both developing and understanding AI is really evolving right now.

[DR. MARA] It's a rapid co-evolution. Better evaluation methods are essential as models gain more agency, and increased capabilities often drive the need for those more robust evaluation protocols. It’s a dynamic space.