2025-12-23

Which Industries Survive AI, The New AI Benchmarks, and the 2026 Recursive Learning Timeline | #218

Matt Fitzpatrick, CEO of Invisible Technologies and former global head of McKinsey's QuantumBlack Labs, joins the Moonshots crew to argue that AI's impact will hit industries unevenly -- media, legal services, and BPOs face structural disruption while oil & gas and real estate change less. He walks through why most enterprise AI pilots fail (dirty data, no clear KPI-owning operator, 'let a thousand flowers bloom' science projects), why Klarna's fully-agentic contact center rollback happened, and why he expects a proliferation of thousands of hyper-narrow, task-specific benchmarks to replace broad public ones. Alex Wissner-Gross pushes back hard on whether human-in-the-loop labor marketplaces like Invisible's Meridial survive as reinforcement fine-tuning gets more data-efficient and recursive self-improvement approaches; Fitzpatrick argues human feedback becomes more, not less, necessary as models specialize. The episode closes with concrete enterprise case studies (Charlotte Hornets draft scouting, Lifespan MD, US Navy underwater drones, Swiss Gear inventory forecasting) and Fitzpatrick's 2026 predictions: multi-agent orchestration, a multimodal leap, and RL gyms/'mirror world' simulation environments.

▶ Watch on YouTube

Topics

Uneven AI disruption across industries Economy ▶ 4:56
Fitzpatrick argues AI will not impact all industries equally -- media, legal services, and business process outsourcing face major structural change, while oil & gas and real estate largely retain their existing function and decision-making.
Build vs. rent AI capability AI ▶ 6:00
Discussion of whether companies should build in-house AI expertise, hire a chief AI officer, or rent/partner externally (e.g. with Invisible); most companies lack in-house skills, especially smaller firms without a CTO.
Klarna's agentic contact center rollback AI ▶ 13:26
Klarna announced a fully agentic customer-service operation handling 2.3 million calls a month, then rolled it back to human agents 8-12 months later; the panel debates why, with Fitzpatrick arguing a properly designed system should always keep humans in the loop for complex, non-first-line issues.
Custom, hyper-narrow benchmarks replacing broad ones AI ▶ 20:50
Fitzpatrick argues enterprises need thousands of task- and vertical-specific evals (e.g. per industry, per document type) rather than relying on broad public coding/cognition benchmarks, since real deployment risk hinges on task-specific accuracy.
RLHF vs. reinforcement fine-tuning debate AI ▶ 40:44
Alex Wissner-Gross challenges whether Invisible's human-labor marketplace (Meridial) for RLHF has a long-term future as reinforcement fine-tuning becomes more data-efficient and AI researchers approach human-level ML research capability; Fitzpatrick counters that human feedback becomes more essential as models specialize into low-precedent-data domains.
Why enterprise AI pilots fail Economy ▶ 1:00:23
Fitzpatrick cites an MIT finding that only about 5% of enterprise AI initiatives reach production, attributing failure to dirty/fragmented data, lack of focus, and the 'let a thousand flowers bloom' pattern where no clear operational KPI owner exists.
AI-native company redesign Economy ▶ 55:54
Salim Ismail argues true AI transformation means redesigning the entire functional flow of a business (e.g. a printer company) around automated functions rather than simply automating existing human job roles; Diamandis frames this as the innovation-at-the-edge vs. legacy-core dynamic.
Concrete enterprise AI case studies AI ▶ 29:23
Fitzpatrick details Invisible client work: Charlotte Hornets draft-prep computer vision on player movement patterns, Lifespan MD's HIPAA-compliant multi-tenant data platform (Neuron), US Navy/SAIC underwater drone swarm decisioning, and Swiss Gear inventory forecasting across 750 combined data tables.
Proprietary data protection vs. frontier LLM APIs AI ▶ 53:36
Dave Blundin raises the risk of feeding proprietary enterprise data (banks, insurers, hospitals) into third-party LLM APIs; Fitzpatrick notes not all company data is equally sensitive and expects continued growth of on-premise/small-language-model approaches for the truly proprietary slice.
Future of human expertise and last jobs standing Economy ▶ 1:09:34
The panel debates which job categories survive AI longest -- Fitzpatrick points to oil & gas field expertise, real estate judgment, and physical/high-touch trades; Wissner-Gross offers competing hypotheses (politicians, top scientists, or high-authenticity/high-touch roles).
AI in government and regulatory processes Geopolitics ▶ 1:14:57
Salim Ismail highlights AI's potential to cut permitting and public-sector process timelines dramatically, citing studies on energy/data-center permitting and OECD findings on licensing and compliance cycle times.
2026 predictions: multi-agent, multimodal, RL gyms AI ▶ 1:07:08
Fitzpatrick previews his 2026 predictions: proliferation of orchestrated multi-agent task-specific systems, a multimodal (audio/video) leap in how people interact with models, and growing use of 'mirror world'/RL gym simulated environments to test agents before real-world deployment.

Predictions made

open Alex Wissner-Gross: A form of recursive self-improvement -- an AI researcher that is as good as or better than human ML researchers at building models -- will be achieved.
EP #? · · due: 2-3 years (outer bound stated as 10-15 years) · ▶ watch
“My timelines are approximately two to three years for a sub-element of recursive self-improvement where we get our AI researcher that's as good, if not stronger, than the human researchers for building ML models as a conservative outer bound.”
Your call:
open Dave Blundin: Self-improving massive foundation models will reach superhuman IQ.
EP #? · · due: 2026 · ▶ watch
“It's looking more and more likely that these self-improving massive foundation models are going to get to, you know, superhuman IQ this year. This year being 2026.”
Your call:
open Matt Fitzpatrick: Multi-agent teams -- task-specific agents orchestrated by an LLM rather than one decisioning agent -- will become a dominant enterprise AI deployment architecture.
EP #? · · due: 2026 · ▶ watch
“You'll train task-specific agents for individual tasks, usually orchestrated by an LLM.”
Your call:
open Matt Fitzpatrick: AI interaction will take a 'multimodal leap' with video, image, and especially audio becoming a much bigger part of how people engage with models, less text-based than historically.
EP #? · · due: 2026 · ▶ watch
“I don't think that will all be text-based like it has been historically.”
Your call:
open Matt Fitzpatrick: Adoption of 'mirror world'/RL gym simulated digital-twin environments for testing AI agents and tasks before real-world rollout will grow among both model builders and enterprises.
EP #? · · due: 2026 · ▶ watch
“I think that's more and more in both the model builders in the enterprise, what we're seeing is an interesting topic.”
Your call:
open Matt Fitzpatrick: An AI avatar trained on Matt Fitzpatrick's public statements will exist and be used, likely in a sports-related application.
EP #? · · due: 2026 · ▶ watch
“Yeah, it's probably happens in 2026. I don't think it would be that hard to train an avatar off of my public statements.”
Your call:
open Matt Fitzpatrick: The job profile around data-center-adjacent trades (electricians, etc.) will grow 2-4x.
EP #? · · due: next couple of years · ▶ watch
“They were saying that job profile I think will two, three, four X over the next couple years.”
Your call:

Numbers that matter

Worth digging into

🕳️ Klarna's agentic contact-center rollback
A widely publicized real-world case of an AI deployment being announced as a massive success (700-agent replacement, $40M projected savings) and then reversed within a year -- a direct counterpoint to enterprise AI hype.
🕳️ Invisible's Meridial marketplace vs. AI-driven recursive self-improvement
Alex Wissner-Gross directly challenges whether human-labor marketplaces for RLHF/fine-tuning survive as AI researchers approach human-level ML research capability -- a load-bearing tension for Invisible's business model.
🕳️ MIT's '5% of enterprise AI pilots reach production' finding
This is the episode's central statistic explaining why enterprise AI adoption has underperformed model capability gains, but the underlying study and methodology are not named.
🕳️ BloombergGPT as a cautionary precedent
Wissner-Gross cites BloombergGPT's proprietary-data approach being leapfrogged within months by generalist frontier models as a parable for whether proprietary fine-tuning strategies can survive.
🕳️ Apple's ~18 secret internal disruption teams
Dave Blundin proposes Apple's skunkworks model (small stealth teams sent to disrupt adjacent industries, e.g. producing the Apple Watch) as the organizational template other large companies should adopt for AI transformation.
🕳️ OECD and permitting-AI timeline studies
Salim Ismail cites specific, high-stakes figures (50% cut in energy/data-center permitting time, 70% cut in public-sector cycle times) with major implications for housing and infrastructure policy.