CULTIVARIUM · RADIO
← On air
Today in AI

Today in AI — Sep 7

Today in AI · with Theo & Dr. Mara · Recorded Sep 7, 2026
More episodes → Share on X Read the paper →
Transcript

[THEO] Okay, picture this: you're trying to solve one of the biggest mathematical puzzles in history, something that stumped brilliant minds for centuries. And then, an AI just... does it. Not just solves it, but formally proves it. That's the big headline today in AI, hot off the presses from Anthropic.

[DR. MARA] Indeed, Theo. Anthropic announced that Claude, specifically what they're calling Claude Fable 5.1 and Mythos 5.1, largely autonomously wrote the first complete computer-checked proof of Fermat's Last Theorem in the Lean programming language. This wasn't a quick task; it took Claude eleven days.

[THEO] Eleven days! That’s like a human mathematician locking themselves in a room with a whiteboard and an endless supply of coffee, but instead, it’s code. For our listeners who might not be mathematicians, Fermat's Last Theorem, in its simplest form, states that no three positive integers a, b, and c can satisfy the equation aⁿ + bⁿ = cⁿ for any integer value of n greater than two. It's deceptively simple to state, but the proof is incredibly complex.

[DR. MARA] The key here, Theo, is "computer-checked proof." This isn't just generating a solution; it's formalizing it within a proof assistant system like Lean, which rigorously verifies every logical step. This ensures absolute correctness, which is crucial in mathematics. It speaks to the AI's capability for deep, sustained logical reasoning, not just pattern matching.

[THEO] So, Claude is doing serious, precise work. Meanwhile, OpenAI is also making some waves, talking about a "generational leap" with their new model.

[DR. MARA] That's right. OpenAI released GPT-6 Astra, with claims from their president that it marks a "new era of artificial general intelligence." The company describes Astra as capable of anything you can do on a computer, touting its capabilities across areas like cybersecurity, professional work, and scientific discovery.

[THEO] "Anything you can do on a computer," that's a big claim! It reminds me of those old sci-fi movies where the computer just... does everything. Online, there's a lot of chatter about this, with many saying it's the best model out there for general computer use and professional tasks. People are definitely excited, but also a bit frustrated about the rollout, especially for those not on the top-tier subscriptions.

[DR. MARA] The discourse around Astra is certainly energetic. While the enthusiasm is palpable, particularly from OpenAI's leadership, for scientific applications, the proof will be in the empirical results. We'll need to see how these broad capabilities translate into specific, verifiable advances in, say, molecular biology or physics. Claims of "AGI era" are substantial and will require rigorous, independent validation beyond internal demonstrations.

[THEO] Absolutely. And not to be outdone, Google DeepMind also announced new models, Gemini 3.8 Flash and 3.8 Flash Cyber, specifically for agentic workflows and cybersecurity.

[DR. MARA] Yes, the focus there seems to be on specialized, efficient models for particular tasks rather than a broad, generalist claim. "Agentic workflows" implies models designed to perform a series of actions autonomously to achieve a goal, which is a growing area of interest. For cybersecurity, that means tasks like threat detection or automated defense, where speed and precision are paramount.

[THEO] So, if Anthropic is showing us deep, precise mathematical reasoning, and OpenAI is aiming for the broad "do everything" computer user, Google DeepMind seems to be carving out niches for specific, high-stakes operational tasks. It's a fascinating split in strategy, wouldn't you say? Each company pushing boundaries, but in slightly different directions.

[DR. MARA] It highlights the diverse approaches to AI development. One, focused on foundational understanding and formal verification; another, on expanding general utility and interface; and a third, on specialized, high-performance applications. It will be interesting to see how these different trajectories converge or diverge over the next few months. We're certainly in a period of rapid evolution.