CULTIVARIUM · RADIO
← On air
Today in AI

Today in AI — Aug 31

Today in AI · with Theo & Dr. Mara · Recorded Aug 31, 2026
More episodes → Share on X Read the paper →
Transcript

[THEO] Alright, Dr. Mara, let's dive into the AI news. What's genuinely new out there this week, beyond the usual chatter?

[DR. MARA] Theo, I think the biggest news—and potentially the most impactful, if it holds—is the buzz around OpenAI's new model, 'Bel.'

[THEO] "Bel," okay. So, picture this: we're talking about the core intelligence that powers something like ChatGPT. It's the brain, the engine. And the headlines are saying "10-trillion-parameter model." That's… a really big number, right? For context, what does "parameters" even mean in this world?

[DR. MARA] Precisely. For listeners who might be more familiar with molecular mechanisms than neural networks, parameters in an AI model are essentially the millions—or in this case, trillions—of numerical values that the model learns during its training. Think of them as the adjustable knobs and switches within the model's architecture. The more parameters, generally, the more complex relationships the model can theoretically learn and represent, and often, the more capable it becomes at understanding and generating data. It’s a measure of its capacity for learning.

[THEO] So, if a previous model was like a moderately complex organism, say, a fruit fly, a 10-trillion-parameter model would be… a blue whale? Just in terms of sheer internal complexity?

[DR. MARA] A blue whale is a good analogy for scale, yes. While the actual functional increase isn't always linear with parameter count, a jump to 10 trillion parameters, if accurate, represents a significant increase in the foundational learning capacity of the model. OpenAI has reportedly completed its pre-training. This means the model has absorbed an enormous amount of data and learned those 10 trillion parameters, setting the stage for its deployment.

[THEO] That's huge. And speaking of deployment, what are the frontier labs doing beyond just making bigger models? Anthropic, for instance, seems to be pushing on how Claude, their large language model, is actually used and evaluated.

[DR. MARA] Yes, Anthropic is making some interesting moves regarding transparency and safety. They've started opening up Claude's usage data to external researchers. This isn't just about sharing, it's about allowing independent analysis of how their AI is being used in real-world conversations. For example, they ran a pilot with groups like Stanford's SALT Lab, where researchers analyzed hundreds of thousands of Claude conversations. The goal here is to understand interaction patterns and identify potential issues or biases that might not be apparent to the developers alone.

[THEO] So, it's like a clinical trial for an AI model, letting outside doctors look at the patient's charts? That seems genuinely useful for understanding how these things behave in the wild.

[DR. MARA] It is. And building on that, Anthropic also reported on Claude-based agents autonomously developing methods to mitigate what they called "alignment failures." This suggests a push towards more self-correcting AI systems, addressing safety challenges directly. Meanwhile, Google DeepMind is trying to standardize how we even *judge* these models, piloting what they call the "world's first double-blind AI evaluations."

[THEO] Double-blind, like in drug trials! That's fascinating. So, neither the AI nor the human evaluator knows who or what they're comparing against? It really tries to cut through the inherent biases.

[DR. MARA] Exactly. In a field often driven by subjective benchmarks and developer-led evaluations, a truly double-blind assessment could bring a new level of rigor to comparing AI system performance and safety.

[THEO] Shifting gears a bit, what's bubbling up in the research papers on arXiv? I'm seeing a lot about robotics and understanding complex physical systems.

[DR. MARA] Indeed. There's a notable trend in robotics, particularly around dexterous manipulation. One paper describes "Aero Hand Open," a tendon-driven robotic hand designed for complex tasks. Tendon-driven hands are interesting because they move the heavy motors away from the joints, making the hand lighter and more agile, which is crucial for fine manipulation.

[THEO] Like a puppeteer pulling strings, right? The power is off-stage, allowing the puppet to be light and expressive.

[DR. MARA] A good analogy. Another paper, "ChainSplat," is tackling the notoriously difficult problem of manipulating deformable linear objects like cables or hoses. They're using a physics-inspired model to learn the dynamics of these objects from video data, which is a major challenge for robots currently. If robots can reliably manipulate flexible items, it opens up a huge range of applications, from manufacturing to disaster relief.

[THEO] And online, beyond the big announcements, what's the general chatter about? There's always some debate swirling around these developments.

[DR. MARA] There's been a lot of discussion about infrastructure. OpenAI leadership has been emphasizing the importance of custom hardware, mentioning their "Jalapeño" custom inference chip. The argument being made is that specialized silicon is crucial for getting more intelligence per watt and achieving faster responses, which directly impacts the user experience and the scalability of these massive models.

[THEO] So, it’s not just about the brain, but also the nervous system and the muscles that make it move efficiently.

[DR. MARA] Precisely. There's also been considerable discussion about cyber defense, with OpenAI and others calling for a global effort to strengthen defenses using AI. The argument is that this is a critical, time-sensitive issue that requires collaborative action, not just competition. And there’s been some discourse around partnerships and access, with OpenAI ending a partnership with Cursor, which some online interpreted as tightening control over model access, especially as companies get acquired. These are the kinds of tensions that emerge as the technology matures and becomes more integrated into industry.