Today in AI — Aug 11
Transcript
[THEO] Welcome back to The Dish! Theo here with Dr. Mara, and we're diving into the latest from the world of AI. Mara, it feels like the big labs are all flexing their math muscles this week.
[DR. MARA] Indeed, Theo. It seems mathematical prowess, or at least the demonstration of it, is a key metric right now. Anthropic and OpenAI have both been showcasing their models' capabilities in this domain.
[THEO] Let's start with Anthropic, because they had a pretty wild claim. They said an unreleased version of Claude took a crack at the Riemann Hypothesis? That's, like, one of the biggest unsolved problems in math. Did it… solve it?
[DR. MARA] Not quite, Theo. The Riemann Hypothesis remains unsolved. However, Anthropic reported that this internal research model, while attempting the Hypothesis, unexpectedly improved the lower bound on the fraction of zeros of the Riemann zeta function from 41.6% to 67.2%.
[THEO] Okay, so it didn't solve the big one, but it made a significant step on a related, very complex problem. For listeners not steeped in number theory, the Riemann Hypothesis is about the distribution of prime numbers, and its proof would have massive implications across mathematics. Finding those "zeros" is a crucial part of understanding it. So, pushing that lower bound, even without a full solution, is still pretty noteworthy. It's like trying to build a rocket to Mars and accidentally inventing a much more efficient engine along the way.
[DR. MARA] A reasonable analogy, Theo. It demonstrates an emergent capability in complex problem-solving, even if the primary objective wasn't met. It speaks to the model's capacity for novel contributions in highly abstract domains. And Anthropic also announced they're making Claude Code's auto mode the default, aiming for even less human oversight in programming tasks.
[THEO] So, Claude's getting smarter at coding *and* doing some surprising math. Meanwhile, OpenAI is teasing their next big thing, Astra. They say it's solved ten long-standing math problems. Are we seeing an AI math-off?
[DR. MARA] It would appear so. OpenAI has revealed Astra, an unreleased model, which they claim has tackled ten previously unsolved mathematical challenges. This signals a similar push into complex logical and mathematical reasoning. However, they've also indicated a pause in some Astra development due to security concerns.
[THEO] That's interesting. So, it's powerful, but they're hitting the brakes? I saw some chatter online about this – the argument seems to be that while these models are becoming incredibly capable, especially in areas like cybersecurity, there's a real need for caution before widespread release. Some are saying it’s important to take more time to ensure safety, even if it delays public access to powerful tools.
[DR. MARA] Precisely. OpenAI has stated that after evaluating Astra, they're treating it as their first "critical" model for cybersecurity under their Preparedness Framework. This implies significant agentic coding advancements, prompting additional safeguards. It highlights the growing tension between rapid deployment and robust safety protocols for increasingly capable AI.
[THEO] Right, so it's not just about what these models can *do*, but what they *could do* if misused. And on a related note, Google DeepMind just announced their WeatherNext AI model has achieved a breakthrough in forecasting cyclones. That's a huge practical application.
[DR. MARA] Indeed. Accurate cyclone forecasting has immense humanitarian and economic implications. If DeepMind's model represents a significant advance, it could lead to more timely and effective disaster preparedness. This moves beyond theoretical math into real-world, life-saving applications.
[THEO] So, we've got AI pushing the boundaries of abstract math, getting better at coding, and potentially saving lives by predicting severe weather. It's a busy week in AI, with a lot of focus on capability, but also, importantly, on caution. Fascinating stuff, Mara.
[DR. MARA] Always, Theo.