# What are we scaling?

Source: https://www.youtube.com/watch?v=_zgnSbu5GqE
Recap page: https://rapidrecap.app/video/_zgnSbu5GqE
Generated: 2025-12-23T21:01:58.656+00:00

---
## Quick Overview

The speaker argues that current AI progress, particularly in scaling reinforcement learning (RL), is misleading because it focuses too much on data progress (like better prompting/data) rather than fundamental algorithmic breakthroughs, suggesting that without true human-like learning capabilities, scaling will eventually hit diminishing returns and that economic diffusion lags in AI adoption are due to missing capabilities, not just slow diffusion.

**Key Points:**
- The speaker believes current AI progress is heavily reliant on data scaling (data progress) rather than algorithmic breakthroughs, making the term 'algorithmic progress' misleading.
- Scaling Reinforcement Learning (RL) requires orders of magnitude more compute (potentially 1,000,000x scale-up) than inference scaling to achieve a similar boost to a GPT-level model.
- Human labor is valuable precisely because it is not 'shleppy' (easy) to train, enabling on-the-job learning and adaptation that current models lack.
- The speaker cites a blog post suggesting that future AGI may rely on continually learning agents that can rapidly internalize skills from experience, unlike current models trained on static datasets.
- The current AI landscape is characterized by an 'economic diffusion lag' because models lack the general, context-specific skills humans possess, leading to slow adoption outside of narrow tasks.
- The speaker anticipates that if the fundamental capability gap isn't closed, the immense investment in scaling current methods will eventually yield diminishing returns, leading to a much slower path to AGI.

![Screenshot at 00:12: The speaker explicitly states that the current approach of training on verifiable outcomes is 'doomed' unless models become closer to human-like learners, highlighting the central critique of scaling laws.](https://ss.rapidrecap.app/screens/_zgnSbu5GqE/00-00-12.jpg)

**Context:** The video features a speaker, likely an AI researcher or commentator, discussing the current state and future trajectory of Artificial Intelligence progress, specifically contrasting the scaling laws observed in large language models (LLMs) with the necessary steps toward Artificial General Intelligence (AGI). The context is a critical analysis of whether simply increasing compute and data (scaling) is sufficient for achieving AGI, or if fundamentally new learning mechanisms are required.

## Detailed Analysis

The speaker argues that the perceived rapid progress in AI, particularly in scaling reinforcement learning (RL) for LLMs, is largely an illusion driven by data progress (better data, better prompting) rather than true algorithmic breakthroughs. He cites evidence suggesting that achieving human-level capability via current methods would require an astronomical 1,000,000x scale-up in total RL compute compared to the 100x scale-up needed for inference scaling to match pre-training gains. This inefficiency is linked to the fact that RL training receives far less information per FLOP than next-token-prediction training. The speaker emphasizes that human workers are valuable precisely because they learn contextually on the job, possessing judgment, situational awareness, and skills that current models lack, which hinders broad economic deployment. He references a blog post suggesting that future AGI will come from continually learning agents that can rapidly assimilate new skills from experience, not just from massive, static pre-training. The speaker concludes that the current focus on scaling without developing these human-like learning capabilities will lead to diminishing returns and that 'economic diffusion lag' is a symptom of missing fundamental capabilities, not just slow technology adoption.

### What are we scaling?

- Implies dissatisfaction with current scaling focus
- Scaling RL requires 1,000,000x compute for GPT-level boost
- Current approach to training on verifiable outcomes is 'doomed'

### Human labor is valuable precisely because it's not shleppy to train

- Humans learn contextually on the job
- AI models lack context-specific skills needed for most jobs
- RL training is inefficient (less info per FLOP than pre-training)

### Economic diffusion lag is cope for missing capabilities

- Models are not near human-like general intelligence
- Current progress is exploitable but not robustly general
- Future progress requires continual learning agents, not just bigger static models

### Goal post shifting is justified

- Previous benchmarks (like GPT-3) were based on pre-training scaling
- Contextual learning capabilities demonstrate power that current models lack

### RL scaling is laundering the prestige of pretraining scaling

- RL compute started small, making gains look huge relative to input
- Scaling current methods won't solve fundamental lack of human-like generalization

![Screenshot at 00:13: Text overlay showing the first point: "01. What are we scaling?"](https://ss.rapidrecap.app/screens/_zgnSbu5GqE/00-00-13.jpg)
![Screenshot at 00:31: Text overlay showing the second point: "02. Human labor is valuable precisely because it's not shleppy to train."](https://ss.rapidrecap.app/screens/_zgnSbu5GqE/00-00-31.jpg)
![Screenshot at 00:45: Text overlay showing the third point: "03. Economic diffusion lag is cope for missing capabilities."](https://ss.rapidrecap.app/screens/_zgnSbu5GqE/00-00-45.jpg)
![Screenshot at 01:07: Text overlay showing the fourth point: "04. Goal post shifting is justified."](https://ss.rapidrecap.app/screens/_zgnSbu5GqE/00-01-07.jpg)
![Screenshot at 02:24: Text overlay showing the fifth point: "05. RL scaling is laundering the prestige of pretraining scaling."](https://ss.rapidrecap.app/screens/_zgnSbu5GqE/00-02-24.jpg)
![Screenshot at 06:37: Text overlay showing the final point: "06. Broadly deployed intelligence explosion."](https://ss.rapidrecap.app/screens/_zgnSbu5GqE/00-06-37.jpg)
