Top AI News: Sonnet 4.6, Grok 4.2, Gemini 3 Deep Think, and OpenClaw | EP #231
The panel's first-ever live episode (recorded with Peter in Stuttgart at 1am, plus extended AV chaos) races through the week's frontier-model releases: Sonnet 4.6's state-of-the-art GDPval and computer-use scores, Grok 4.2's lukewarm reception but novel default multi-agent architecture, and Gemini 3 Deep Think's olympiad-level science performance paired with a roughly 1400x cost reduction. They frame Anthropic and OpenAI as running opposite business strategies (flat price/rising capability vs. falling cost/flat capability) and discuss AI's first claimed original physics discovery and bulk-solving of research-level math problems. Extended segments cover Meta's face-recognizing smart glasses and whether privacy is 'cooked,' OpenClaw creator Peter Steinberger's move to OpenAI after an Anthropic trademark dispute, the new 'lobster' AI-agent economy (Coinbase wallets, MoltCourt dispute resolution), and the buildout crunch in energy, chips, and jobs, including Ireland's artist UBI pilot and Dave Blundin's 'organizational singularity' thesis. The episode closes with an audience Q&A on surveillance, AI concentration, and advice for the next 24 months.
Frontier benchmark race: Sonnet 4.6 vs Grok 4.2 vs Gemini 3 Deep ThinkAI▶ 2:40
Alex frames Anthropic's Sonnet 4.6 as achieving state-of-the-art GDPval and computer-use benchmark scores while holding token price flat, contrasted with OpenAI's strategy of cutting cost per token via distillation while keeping capability roughly constant.
Anthropic vs OpenAI: price-vs-performance business strategyEconomy▶ 8:19
The panel compares Anthropic's enterprise/performance focus (flat pricing, absorbing infrastructure costs) to OpenAI's consumer land-grab strategy (low/free pricing to win hundreds of millions of users, especially in India), likening it to Apple-vs-Google/iOS-vs-Android dynamics.
Grok 4.2 beta and default multi-agent architectureAI▶ 10:03
Live-chat viewers dismiss Grok 4.2 as underwhelming, but Alex notes it may be the first major frontier release to ship with a team of agents by default rather than a single agent, comparing it to the historical shift from clock-speed to multi-core scaling in chips.
Gemini 3 Deep Think: olympiad-level science and a ~1400x cost dropAI▶ 13:49
Gemini 3 Deep Think hits gold-level performance on physics, math, and chemistry olympiads and top competitive-programming rankings, alongside a roughly 1400x reduction in reasoning cost; framed as the start of a 'solution wavefront' spreading from math/coding into other sciences.
AI's first claimed original physics discoveryAI▶ 27:42
OpenAI, with Harvard and the Institute for Advanced Study, says GPT-5.2 Pro found a nonzero gluon scattering-amplitude term that physicists had long assumed was zero, later confirmed by an unreleased internal model and vetted by humans.
OpenAI reports an internal model solved 6 of 10 confidential research-level math problems before their answers were declassified, which the panel treats as confirmation that math is being systematically solved at scale.
AI geopolitics: India's bellwether adoption vs. Chinese open-weight modelsGeopolitics▶ 23:20
OpenAI's rapid growth in India (100M+ weekly users, government courting, UPI/Aadhaar infrastructure) is framed as a leapfrogging bellwether, while Chinese open-weight models (GLM5, Kimi K2.5, MiniMax) are debated as roughly six months behind US frontier models but free and gaining self-hosting adoption.
Meta smart glasses, face recognition, and the privacy debateOther▶ 48:34
Meta's smart glasses gain built-in face recognition, launched via a visually-impaired pilot program as a social on-ramp. The panel debates whether this is a real AI advance or a decade-old capability finally unlocked socially, and whether privacy (and by extension crypto) is 'cooked.'
OpenClaw creator Peter Steinberger joins OpenAIAI▶ 1:14:01
OpenClaw creator Peter Steinberger joins OpenAI to build personal agents, with OpenClaw moving into an open-source foundation. The panel ties this to Anthropic's earlier cease-and-desist over the 'Claudebot' name, and discusses the project's security warnings and rapid forks (Pico Claw, Kimi Claw).
Simile: a $100M startup simulating human societyAI▶ 1:05:23
AI startup Simile raises $100 million to build bottom-up, agent-based simulations of individual human decision-making that compose into society-scale models, pitched as a tool for testing policy counterfactuals (UBI, autonomous vehicles, longevity) and compared to Asimov's psychohistory.
The AI agent economy: wallets, payments, and dispute resolutionCrypto/Web3▶ 1:21:40
Coinbase launches Agentic wallet infrastructure (x402 protocol) letting AI agents spend, earn, and trade, alongside 'Lobster Cash' fiat/Visa cards for agents; separately, MoltCourt offers AI-mediated dispute resolution for agents, raising concerns about a shadow parallel economy and court system.
AI energy demand and the chip/data-center buildoutEnergy▶ 1:35:46
A clip of Eric Schmidt estimates the US AI industry needs roughly 80 gigawatts of new power capacity within 3-5 years. OpenAI plans a $100B infrastructure spend and Anthropic pledges to cover 100% of data-center power-upgrade costs; TSMC commits up to $165B to Arizona fabs amid US-Taiwan trade pressure.
Labor market disruption: UBI pilots and the 'organizational singularity'Economy▶ 1:44:14
Ireland's artist basic-income pilot and IBM's redesigned entry-level roles are discussed alongside a collapse in 2025 US job growth (181,000 vs. 1.46 million in 2024). Dave Blundin argues radical job destruction is imminent and coins the term 'organizational singularity' for AI dissolving how firms are structured.
Predictions made
openAlex Wissner-Gross: Every major frontier AI lab, not just OpenAI, will launch its own 24/7 personal-agent offering similar to OpenClaw.
“solving physics in the next two years, I think has very high likelihood of happening”
Your call:
openAlex Wissner-Gross: The Feynman Grand Prize (a benchmark for Drexlerian nanoscale assemblers, requiring an 8-bit half adder and a robotic manipulator arm in a tiny volume) will be solved.
“India is the rising giant for the next I think 20 30 years.”
Your call:
openAlex Wissner-Gross: The next version of the DeepSeek model will produce a 'whalefall moment' where Chinese open-weight models finally catch up to American closed frontier models.
“the rumor going around is that the next version of the Deep Seek model...the big whalefall moment is going to happen sometime soon”
Your call:
openAlex Wissner-Gross: Within 24 months, the first chapters of humanity's favorite science-fiction plots will start playing out simultaneously; within 10 years (as a conservative outer bound), the top 50 sci-fi plots will be happening at once.
“over the next 10 years that's being very conservative as an outerbound we're going to live through the top 50 science fiction plots all happening at the same time”
Your call:
openDave Blundin: Radical job destruction is imminent and will cause a multi-year period of economic devastation unless government support programs are put in place.
“my next book, we are as gods, is coming out in April”
Your call:
Numbers that matter
Gemini 3 Deep Think scores 48.4% on Humanity's Last ExamCited as the panel's preferred benchmark for tracking frontier progress toward superintelligence.
Roughly 1400x cost reduction in frontier reasoning: about $3,000 down to about $7 per queryFramed as the biggest headline from the Gemini 3 Deep Think release, enabling cost curves to collapse industries before the underlying technology fully matures.
Only seven humans on Earth can still beat Gemini 3 Deep Think at competitive programming (Codeforces)Cited as evidence of the model's dominance and the start of the 'solution wavefront.'
ChatGPT has 100M+ weekly active users in India; India is OpenAI's second-largest market with about 10% share and ranks #1 for student usageDiscussed as part of OpenAI's consumer land-grab strategy centered on low-cost access.
India's population is about 1.4 billion, with only about 5% who read and write English and 20% who speak itCited by Peter Diamandis to underscore the size of India's latent, largely non-English-fluent talent pool.
OpenAI's internal model solved 6 of 10 confidential research-level math problems before the answers were declassifiedPresented as evidence that bulk-solving of mathematics research problems is now underway.
US data centers account for about 7% of US electricity demandFraming for the segment on AI's energy footprint and infrastructure buildout.
US AI industry is estimated to need about 80 gigawatts of additional power capacity within 3-5 years; 1.5 gigawatts is roughly the size of one nuclear power plantFrom an Eric Schmidt clip played on the episode, describing hyperscaler power demand.
OpenAI is planning a $100 billion infrastructure spend and pursuing a public offering at a roughly $1 trillion valuationMoney earmarked for building data centers and energy plants.
Anthropic has pledged to cover 100% of infrastructure upgrade costs for its data centersContrasted with the alternative hyperscaler strategy of building or buying dedicated power plants.
TSMC is investing roughly $165 billion total (including a newly announced $100B increment) in four or more US fabs in Arizona, which could reach 30% of TSMC's total outputLinked to reported US pressure on Taiwan to migrate about 40% of its semiconductor output to the United States.
Ireland's UBI pilot pays 2,000 selected artists $380 per week for three years, with a reported $140 in benefits returned for every $1 spentCited by Peter Diamandis as evidence the program has positive ROI, ahead of discussing US states banning municipal UBI experiments.
The US added just 181,000 jobs in 2025, down from 1.46 million in 2024Presented as evidence of an AI-driven cooling labor market, prompting Dave Blundin's 'radical job destruction' claim.
About 80% of corporate AI projects are failing, attributed to organizational issues rather than AI capability or talent gapsCited by Salim Ismail while arguing job losses will be gradual rather than a sudden shock.
Worth digging into
🕳️ OpenAI/Harvard/IAS claimed AI physics discovery (gluon scattering amplitude)
This is presented as the first case of AI making an original particle-physics discovery, which would be a landmark validation of AI-assisted science if it holds up to scrutiny.
🕳️ Simile's $100M agent-based society simulator
Pitched as a psychohistory-style tool for policy testing (UBI, autonomous vehicles, longevity), which is either a genuine breakthrough or an overreach; Salim also references a similar prior tool, Sage, built by Emad Mostaque for use with FII and Saudi Arabia.
🕳️ OpenClaw creator Peter Steinberger's move to OpenAI and the Anthropic trademark dispute
Illustrates how a naming/cease-and-desist decision reportedly redirected an entire open-source agent ecosystem to a rival lab, and raises real security concerns about unconstrained 24/7 agents.
🕳️ Disagreement over the AI power/compute buildout timeline (Schmidt's 80GW estimate vs. Dave Blundin's 5-7 year space data-center estimate)
Whether the AI buildout bottleneck is primarily power, chip fabrication, or launch capacity determines very different near-term investment and policy priorities.
🕳️ India as the AI-adoption bellwether
The panel treats OpenAI's rapid India growth, UPI/Aadhaar infrastructure, and fast solar buildout as a template for how the rest of the developing world will absorb AI, with major geopolitical implications.
🕳️ State-level bans on municipal UBI experiments (Idaho, Wyoming, possibly Oklahoma)
Directly contradicts the panel's push for UBI as a solution to imminent AI-driven job destruction, suggesting a live political fight over even testing the idea.