# Why We Need New AI Benchmarks, Which Industries Survive AI, and Recursive Learning Timelines | #218

Source: https://www.youtube.com/watch?v=7GFKB0oKd9A
Recap page: https://rapidrecap.app/video/7GFKB0oKd9A
Generated: 2025-12-23T17:00:47.371+00:00

---
## Quick Overview

The largest disruption from companies failing to adopt AI will occur in 2026, necessitating that companies focus on identifying 2-3 needle-moving use cases and begin with external vendor RFPs tied to measurable results rather than attempting broad internal overhauls, while the need for thousands of narrow, task-specific benchmarks will supersede broad public ones.

**Key Points:**
- Matt Fitzpatrick predicts the "largest disruption ever in 2026 from companies that don't make this change," stating that "knowledge work as we currently know it" is cooked.
- Companies should not start with a 'let a thousand flowers bloom' approach but must identify 2-3 use cases that "materially move the needle" for the business.
- The initial AI deployment for a key use case should be done via an RFP to a third-party vendor compensated based on results to limit risk, as in-house teams often lack executive experience in this paradigm.
- The focus of benchmarking must shift from broad public metrics (like coding) to "custom evals on highly specific topics" to achieve task accuracy or human equivalence, necessitating thousands of narrow benchmarks.
- The adoption curve for AI in areas like contact centers has been slow due to challenges in handling complex, non-first-line resolution topics and the inherent human preference for talking to other humans.
- Invisible Technologies focuses on tailoring existing LLMs by fine-tuning them on company-specific information, asserting that human-in-the-loop validation (RLHF) remains crucial, especially for reasoning tasks, even as models advance.
- A major challenge for AI implementation is the lack of focus on data as the starting point; companies must focus on the exact data needed for a specific use case, especially unstructured data like text and images, rather than trying to master the entire data repository.

**Context:** This episode of Moonshots features Matt Fitzpatrick, CEO of Invisible Technologies and former Global Head of Quantum Black Labs at McKinsey, discussing the urgent need for enterprise transformation due to AI capabilities. The conversation centers on the impending massive disruption expected by 2026 for companies that fail to pivot to become 'AI companies,' the differing impacts AI will have across industries, and the critical role of specialized data and benchmarking in successful implementation.

## Detailed Analysis

Matt Fitzpatrick asserts that enterprises are moving too slowly relative to AI capabilities, predicting a massive disruption in 2026 for those lagging, specifically noting that knowledge work structure will fundamentally change in sectors like media and legal services, though heavy industries like oil and gas will see less structural change. He advises CEOs facing board inquiries to prioritize identifying 2-3 high-value use cases, pilot one quickly, and ideally use an external vendor compensated by results to manage implementation risk, contrasting this with the traditional ML paradigm of long build times. A major theme is the evolution of benchmarking: broad public benchmarks are insufficient, and companies must develop thousands of "hyperspecific benchmarks" or custom evaluations for every targeted labor category or task to validate performance beyond 80% accuracy. Fitzpatrick also highlights that data quality is the primary failure point, urging companies to focus only on the exact structured and unstructured data required for a specific use case, rather than attempting a massive, multi-year data clean-up. Regarding implementation, he notes that many companies will need to 'rent' expertise externally rather than build the capability entirely in-house. Finally, the discussion touches on the ongoing necessity of human-in-the-loop processes (RLHF), even as models become more powerful, arguing that complex reasoning tasks and company-specific trade secrets require continuous human validation, countering the notion that full autonomy will eliminate the need for human feedback soon.

### 2026 AI Disruption Forecast

- Largest disruption ever expected in 2026 from companies that do not change
- Knowledge work structure faces fundamental change in sectors like legal and BPO
- Oil and gas and real estate will see less structural change.

### Enterprise AI Implementation Strategy

- Focus on 2-3 needle-moving use cases, not 'letting a thousand flowers bloom'
- Start with a pilot, preferably outsourced via an RFP tied to results
- Companies must honestly assess whether to 'buy or rent' AI expertise.

### The Benchmark Problem

- Public benchmarks are useful for model improvement but insufficient for enterprise tasks
- Need for "thousands of new narrow benchmarks" to capture accuracy on specific tasks
- Businesses must become comfortable with custom evaluations ('evals') for modernization tasks.

### Data Foundation for AI

- Lack of focus on data is a primary failure point
- Focus on the exact data needed for the use case, not cleaning all enterprise data
- Most important data for GenAI is often non-system-of-record, unstructured data (images, text).

### Case Studies in Custom AI

- Invisible worked with the Charlotte Hornets using fine-tuned computer vision models to analyze player spatial movement patterns for draft selection
- Lifespan MD uses data consolidation on a HIPPA-compliant platform to create control towers for practice performance and patient outcomes, heavily focusing on automating administrative tasks.

### The Role of Human Feedback (RLHF)

- Human feedback remains necessary as models move into highly specific reasoning tasks
- The idea that autonomous agents will operate with no human in the loop is a 'red herring' for the enterprise for a long time
- The nature of RLHF is changing towards RL gyms and controlled environments.

