Mira Murati's 975B Open Model, Ramin Hasani on Post-Transformer AI, and Demis' AI FINRA | EP #271
The Moonshots crew debates a wave of frontier-lab CEOs (Altman, Musk, Hassabis) calling for AI regulation modeled on FINRA, and a reported White House proposal to peg US open-weight model releases to China's pace, with the group split between seeing this as necessary safety infrastructure and as regulatory capture. Mira Murati's Thinking Machines Lab ships Inkling, a 975B-parameter open-weight MoE model, sparking discussion of on-prem fine-tuning as the new competitive battleground. A London startup's claimed recursive self-improvement result draws sharp pushback from guest Ramin Hasani (Liquid AI), who argues it isn't true weight-level RSI and walks through Liquid's own worm-brain-inspired liquid neural network architecture, its automated architecture-search pipeline (AFMD), and its on-device deployments with Mercedes-Benz and Shopify. The episode closes with Palmer Luckey's controversial call to expand classified national-security patents, AI models now beating physicians on medical benchmarks, and a new enzyme from Revel Pharmaceuticals that reverses age-related protein damage (AGEs) in human tissue.
Frontier lab CEOs call for AI regulation (FINRA-style standards body)AI▶ 4:24
Sam Altman, Elon Musk, and Demis Hassabis have all separately called for AI regulatory bodies, with Hassabis proposing a FINRA-modeled, industry-funded frontier AI standards body operational by end of year. The group debates whether this is genuine safety planning or regulatory capture that locks out smaller labs.
White House capability-ceiling proposal tied to China's open modelsGeopolitics▶ 21:26
A reported White House framework would let US labs release open models as long as they stay at or below the capability level of China's best open-weight model, effectively letting Beijing set the ceiling. The panel calls this a perverse incentive and a 'trade dispute' framing of the AI race.
Mira Murati's Thinking Machines Lab ships Inkling, a 975B open-weight modelAI▶ 26:00
Thinking Machines Lab released Inkling, a 975B-parameter (41B active) mixture-of-experts open-weight model trained on 45 trillion multimodal tokens, positioned as a western counterweight to DeepSeek and Qwen. It benchmarks above Nvidia's Nemotron but below China's GLM 5.2.
Fine-tuning and on-prem model customization as the new business battlegroundAI▶ 35:55
The panel discusses why enterprises increasingly want fine-tuned, on-prem open-weight models rather than sending proprietary data to closed frontier APIs (the 'Alex Karp' data-leakage argument), and debates whether Thinking Machines' fine-tuning-as-a-service bet is durable or a wager on a passing reinforcement-fine-tuning paradigm.
London startup WeCo AI (transcribed as 'Wo AI'/'WICO') claims early experimental evidence of recursive self-improvement, with an outer AI agent rewriting an inner agent's code and research strategy, claiming 8 days of machine self-improvement beat two years of expert human effort. They also publish a 0-3 RSI maturity scale.
Ramin Hasani argues WeCo AI's result is impressive engineering but not real RSI, since it only rewrites code/prompts around fixed-weight models rather than retraining weights, and estimates the compute cost of true weight-level self-improvement is currently intractable.
Malaysian PM's AI digital twin and the rise of leader avatarsAI▶ 1:02:36
Malaysia's PM Anwar Ibrahim is preparing an AI clone of himself for public communications across Malaysia's 135 languages, following Albania's AI cabinet minister. The panel predicts corporate CEOs, religious leaders, and eventually AI Peter/AI Dave-style avatars will follow, potentially becoming bidirectional rather than broadcast-only.
Liquid AI's origin story and liquid neural networksAI▶ 1:14:32
Ramin Hasani recounts founding Liquid AI out of MIT CSAIL research on C. elegans (a 302-neuron worm) to build continuous-time, analog-inspired neural architectures as an alternative to the transformer, and describes Liquid's automated architecture-search framework (AFMD/STAR) for finding efficient architectures per deployment target.
Mercedes-Benz on-device small language model deploymentAI▶ 1:23:08
Liquid AI is shipping a sub-1GB multimodal model that runs on cheap in-car chips ($60-class, 2-8GB RAM), giving Mercedes vehicles offline, private, function-calling access to 700-1,200 in-car functions, rolling out via a 600MB over-the-air update to North American 2022+ vehicles.
Palmer Luckey's push to expand classified national-security patentsGeopolitics▶ 1:39:36
Anduril's Palmer Luckey argues the patent system is a national security liability because disclosure requirements let adversaries harvest US inventions, and proposes massively expanding the 1951 Invention Secrecy Act's classified-patent mechanism. Alex Wissner-Gross calls this a misreading of how the Act actually works and a bad idea overall.
AI diagnostics now beating physicians (GPT-5.6, Meta Muse Spark 1.1)Health▶ 1:48:10
GPT-5.6 set a new high on OpenAI's HealthBench Professional and beat specialty-matched physicians with full internet access in a ~20,000-judgment blind test. Meta's free Muse Spark 1.1 then beat GPT-5.6 on the same benchmark at 7x lower cost, reaching Meta's 3.56 billion daily active users.
Revel Pharmaceuticals (with Calico) published an engineered enzyme called CMLA that reverses advanced glycation end-products (AGEs) - sugar-protein cross-links long considered irreversible drivers of aging - restoring healthy protein structure in human tissue samples from elderly donors.
Predictions made
openPeter Diamandis: A formal AI regulatory body or structure will emerge
“you're going to see like unbelievably kind of models like probably in the next 2 years or so, you know, like models that are like going above our our understanding”
Your call:
openDave Blundin: AI digital-twin avatars of political and public figures will become the dominant medium for political outreach
EP #? · · due: next election cycle, about 2 years · ▶ watch
“it should be easily dominant two years from now in the next election... once people realize that there's no going back. It's going to be huge.”
Your call:
openDave Blundin: The US will trade-embargo countries that don't respect intellectual property rights amid an accelerating AI-driven innovation boom
EP #? · · due: within the next couple of years · ▶ watch
“the likely outcome of that is the US will trade embargo anybody who doesn't respect intellectual property rights... that's going to happen soon like in the next couple of years”
Your call:
Numbers that matter
975 billion total parameters, 41 billion activeSize and mixture-of-experts activation pattern of Thinking Machines Lab's Inkling open-weight model.
45 trillion training tokensInkling was trained on 45 trillion tokens of text, image, audio, and video.
~7 month average lagReported average gap between Chinese and US open-weight model release timelines, cited as the basis for the White House capability-ceiling proposal.
~600,000 patent applications/year, ~323,000 granted (2025), up 40% over 5 yearsUS Patent Office volume cited in the Palmer Luckey patent-reform discussion; the increase is attributed partly to AI-assisted filing.
~6,000 active secrecy ordersNumber of classified patents currently active under the US Invention Secrecy Act of 1951, which Palmer Luckey wants massively expanded.
8 days of machine self-improvement claimed to beat 2 years of expert human effortWeCo AI's self-reported claim for its AIDE² recursive self-improvement system.
Chinchilla ratio of ~20 tokens per parameterCited scaling-law rule of thumb for compute-optimal training token budgets relative to model size.
~350 years of compute to fine-tune a 2B-parameter model under WeCo's frameworkRamin Hasani's illustrative calculation showing the computational intractability of true weight-level recursive self-improvement at current framework scale.
Sub-1GB model size, running on ~$60 chips with 2-8GB RAMLiquid AI's on-device multimodal model footprint for Mercedes-Benz infotainment chips.
600MB OTA update, with future incremental updates as small as ~20MBSize of the over-the-air update bringing Liquid AI's model to North American Mercedes-Benz vehicles from model year 2022 onward.
700-1,200 in-car functionsNumber of vehicle functions accessible via function-calling from Liquid AI's on-device model in Mercedes vehicles.
~1 billion requests, hundreds of millions of users, 10 billion productsScale of Liquid AI's foundation model deployment inside Shopify, in production for about 6 months.
302 neurons, 78% genome similarity to humans, 4 Nobel prizesFacts about C. elegans cited as the biological basis for Liquid AI's original liquid neural network research.
~20,000 physician judgments in blind testScale of the blind comparison in which GPT-5.6's medical answers outperformed specialty-matched physicians with unlimited web access and time.
7x cheaper than GPT-5.6Meta's Muse Spark 1.1 reportedly beat GPT-5.6 on OpenAI's HealthBench Professional (525 clinical tasks) at 7 times lower cost.
It's self-reported, unconfirmed by outside parties, and Ramin Hasani directly disputes that it constitutes true weight-level RSI - a live methodological dispute worth tracking for independent replication.
🕳️ Liquid AI's automated foundation model design (AFMD/STAR) architecture search
An unbiased search process reportedly reconverged on the same gating mechanism from Liquid's original liquid-neural-network paper, suggesting a real architectural signal beyond marketing - relevant to the broader post-transformer architecture debate.
A category of aging damage long considered permanent (AGE cross-linking) was shown reversible in human tissue samples - directly relevant to the Longevity Escape Velocity thesis discussed on the show.
🕳️ White House capability-ceiling proposal pegging US open releases to China's pace
A genuinely novel and contested regulatory mechanism with game-theoretic implications (incentivizing China to advance faster) that the panel found 'perverse' and worth continued scrutiny.
🕳️ Palmer Luckey's proposal to expand classified national-security patents
Directly disputed on air (Alex Wissner-Gross calls it a misunderstanding of the 1951 Invention Secrecy Act), with big implications for how much transformative tech might already be classified/confiscated without public knowledge.