What are we scaling?

Quick Overview

The speaker argues that current AI progress, particularly in scaling reinforcement learning (RL), is misleading because it focuses too much on data progress (like better prompting/data) rather than fundamental algorithmic breakthroughs, suggesting that without true human-like learning capabilities, scaling will eventually hit diminishing returns and that economic diffusion lags in AI adoption are due to missing capabilities, not just slow diffusion.

Key Points: The speaker believes current AI progress is heavily reliant on data scaling (data progress) rather than algorithmic breakthroughs, making the term 'algorithmic progress' misleading. Scaling Reinforcement Learning (RL) requires orders of magnitude more compute (potentially 1,000,000x scale-up) than inference scaling to achieve a similar boost to a GPT-level model. Human labor is valuable precisely because it is not 'shleppy' (easy) to train, enabling on-the-job learning and adaptation that current models lack. The speaker cites a blog post suggesting that future AGI may rely on continually learning agents that can rapidly internalize skills from experience, unlike current models trained on static datasets. The current AI landscape is characterized by an 'economic diffusion lag' because models lack the general, context-specific skills humans possess, leading to slow adoption outside of narrow tasks. The speaker anticipates that if the fundamental capability gap isn't closed, the immense investment in scaling current methods will eventually yield diminishing returns, leading to a much slower path to AGI.

Context: The video features a speaker, likely an AI researcher or commentator, discussing the current state and future trajectory of Artificial Intelligence progress, specifically contrasting the scaling laws observed in large language models (LLMs) with the necessary steps toward Artificial General Intelligence (AGI). The context is a critical analysis of whether simply increasing compute and data (scaling) is sufficient for achieving AGI, or if fundamentally new learning mechanisms are required.

Raw markdown version of this recap