The Perils of the AI Exponential

Quick Overview

The video analyzes the recent deceleration in the improvement rate of AI agent capabilities, specifically citing METR's "Moore's Law for AI agents" chart which shows task length doubling every 7 months, and contrasting this with recent, seemingly slower progress seen in models like GPT-5.3-Codex and Claude Opus 4.6, which exhibit doubling times closer to 1.5 and 4.5 months respectively, suggesting that the rapid exponential gains seen previously are beginning to plateau or that new scaling laws are emerging, which is further supported by commentary from experts like Daniel Rein and Guy Berger questioning the sustainability and methodology of these metrics.

Key Points: METR's "Moore's Law for AI agents" research previously indicated that the length of tasks AIs could double roughly every 7 months (0:05). Recent models like GPT-5.3-Codex (05:20) and Claude Opus 4.6 (07:07) showed faster improvements, with doubling times of 6.5 hours/17 hours and 14.5 hours, respectively, which translates to significantly shorter doubling times than the original 7-month trend (05:00, 05:39). Expert commentary, such as from Daniel Rein (06:50), noted that the measurement is extremely noisy and that the recent apparent jump in coding capabilities might be an outlier due to changes in scaffolding/format (06:18). Economist Guy Berger (10:18) questioned the sustainability of the trend and pointed out that if current growth continues, AI could consume the entire economy by 2028 (07:50), leading to massive unemployment. The Citrini Research memo, "The 2028 Global Intelligence Crisis" (07:25), models this scenario of rapid AI advancement leading to economic collapse and mass joblessness. The discussion concluded that while the overall trend is still positive, there is evidence of saturation in certain areas, and the consistency of the measurement itself is being questioned by experts (06:33, 08:27).

Context: This AI Daily Brief discusses recent developments and debates surrounding the progress rate of AI agents, focusing heavily on the findings and implications of research conducted by METR (Model Evaluation & Threat Research). The core of the discussion revolves around METR's chart tracking AI performance on long tasks over time, often referred to as "Moore's Law for AI agents," and whether the latest flagship models (like GPT-5.3-Codex and Claude Opus 4.6) confirm or challenge the established exponential growth rate.

Raw markdown version of this recap