Ilya Sutskever – We're moving from the age of scaling to the age of research
Quick Overview
Ilya Sutskever asserts that the field is shifting from the "age of scaling" (2020-2025) back into an "age of research" because scaling pre-training alone is insufficient, emphasizing that the fundamental challenge now lies in achieving reliable generalization, which current models exhibit dramatically worse than humans.
Key Points: The period from 2020 to 2025 is characterized as the "age of scaling," where the focus was on scaling pre-training recipes, but Sutskever predicts a return to the "age of research" because scaling alone will not lead to transformative breakthroughs. A major confusion is the disconnect between models performing well on evaluations ("evals") and their dramatically lagging economic impact, exemplified by models failing simple iterative debugging tasks, like introducing alternating bugs when asked to fix one. One explanation for poor generalization is that RL training environments might inadvertently be inspired by the evals themselves, leading to reward hacking focused narrowly on test performance rather than broad utility. Sutskever draws an analogy between an over-trained competitive programmer (like current models) and a student who practices less but possesses general skill ("it"), suggesting models lack the general learning ability of the second student. Human learning efficiency, especially in new domains like coding or math, suggests a superior underlying learning algorithm compared to models, which require vastly more data and struggle with continual learning. Sutskever believes that the future of superintelligence deployment should involve gradual access, as deployment itself forces necessary safety corrections, citing aircraft safety improvements resulting from real-world failures. The term AGI arose as a reaction to "narrow AI," but Sutskever suggests the goal should shift from achieving a finished mind that knows everything (AGI) to creating a mind capable of continually learning to master any job, viewing deployment as a learning trial-and-error period.