# Turing CEO Jonathan Siddharth: Who Wins in Data Labelling & Why 99% of Knowledge Work Will Disappear

Source: https://www.youtube.com/watch?v=yQLOicn2vPU
Recap page: https://rapidrecap.app/video/yQLOicn2vPU
Generated: 2025-12-01T15:38:05.146+00:00

---
## Quick Overview

Jonathan Siddharth, CEO of Turing, asserts that the era of simple data labeling companies is over, replaced by the era of "research accelerators" focused on training superintelligence through complex data generation, particularly via Reinforcement Learning (RL) environments, predicting that 99% of knowledge work will eventually disappear.

**Key Points:**
- Turing's focus has shifted from being a talent marketplace to a research accelerator that trains superintelligence by supplying the critical data pillar needed by frontier labs.
- The required data has shifted from simple instructions (e.g., "write a Python program to sort numbers") to complex, vertically specific tasks that require expert humans to generate data for agentic models.
- Training agents requires teaching tool use, often dominated by Reinforcement Learning where Turing creates RL environments—mini world models tracking system state, input prompts, and output verifiers—for massive scale workflow training.
- Siddharth believes the AI takeoff is a slow, steady process, stating the industry is still in "innings one" regarding the acquisition of vertically focused workflow data.
- The need for custom models remains a permanent requirement for enterprises like insurance companies where smaller, fine-tuned models (e.g., 0.5B to 10B parameters) are faster and more accurate for specific tasks like underwriting, and where data privacy is paramount.
- Siddharth predicts all knowledge work involving looking at a computer, analyzing screens, and using tools will be automated within a 10 to 20-year timeframe, leading to 100x productivity gains for those who adopt the technology.
- The primary moat in the future will be data-driven feedback loops derived from deploying models to touch reality in the enterprise, requiring significant first-mile shape (data acquisition/structuring) and last-mile shape (workflow integration/handholding).

**Context:** Jonathan Siddharth, CEO of Turing, a company scaled to over $350 million in ARR, discusses the evolution of AI development and the role of data providers with host Harry Stebbings. Turing positions itself as a research accelerator supporting frontier AI labs (like OpenAI, Anthropic, DeepMind) by supplying the necessary data to achieve superintelligence, distinct from traditional talent marketplaces or simple data labeling services.

## Detailed Analysis

Siddharth argues that the AI landscape has fundamentally changed, moving away from simple data labeling to complex data generation required for increasingly sophisticated models that are transitioning from chatbots to agentic systems capable of executing multi-step workflows. This transition necessitates training models on tool use via Reinforcement Learning (RL), for which Turing builds massive-scale RL environments simulating business workflows across every industry, function, and role—a scope he estimates as $30 trillion worth of knowledge work. He posits that incumbents failing to adopt these tools will face extinction by agile startups, though he anticipates a slow takeoff for AGI. Furthermore, he details the permanent need for custom, fine-tuned models within enterprises (e.g., insurance underwriting) where smaller models can leverage proprietary historical data without sharing it with frontier labs. The future moat lies in data-driven feedback loops established through real-world deployment, which involves overcoming "first mile shape" (data clean-up) and building workflows for partial autonomy, like the 'cursor for X' concept, although simpler roles like customer support may skip this intermediate step entirely. Siddharth dismisses the notion of an AI bubble, citing the immense capability overhang of current models that are vastly underutilized without proper agentic scaffolding, and predicts that access to intelligence via API will democratize entrepreneurship rather than widen societal gaps.

### Turing's New Mandate

- Shifting from talent marketplace to research accelerator
- Supplying the data pillar for superintelligence alongside research and compute
- Data needs are now complex, requiring expert human input for agentic training.

### The Shift to Agentic Training

- Transitioning from SFT/RLHF for chatbots to RL for agents that execute multi-step workflows
- Agents require training on tool use and calling functions, often built within simulated RL environments.

### Data Generation at Scale

- Turing creates RL environments for every workflow, role, function, and industry ($30 trillion knowledge work scope)
- This synthetic data creation is analogous to AlphaZero mastering Go through self-play.

### Enterprise Customization vs. General Models

- Permanent requirement for smaller, fine-tuned models on-prem for enterprises needing proprietary knowledge distillation (e.g., insurance underwriting)
- Frontier labs may not want proprietary data used by competitors.

### The Future of Knowledge Work

- Prediction that all computer-based knowledge work disappears within 10-20 years, leading to 100x human productivity
- This abundance of cheap intelligence should boost entrepreneurship by lowering capital constraints for non-technical founders.

### Moats and Competitive Advantage

- Future moats rely on data-driven feedback loops from real-world deployment, requiring solving first-mile shape (data structuring) and last-mile shape (workflow integration/handholding).

### The SAS Industry Fate

- Siddharth believes SAS as known is over because AI applications are easier to build, creating risks from companies building themselves or being 'sonic boomed' by foundation model companies moving into the app layer.

