Rich Sutton, The OaK Architecture: A Vision of SuperIntelligence from Experience - RLC 2025

Quick Overview

Rich Sutton presents the OAK architecture as a vision for achieving superintelligence through experience, emphasizing that true AI progress lies in improving reinforcement learning algorithms rather than relying on non-experiential methods like large language models. OAK is designed to be domain-general, experiential, and open-ended in sophistication, learning entirely at runtime from the complex, big world rather than through pre-programmed design-time knowledge. The architecture aims to enable agents to discover high-level abstractions and reasoning capabilities through self-generated sub-problems, mirroring natural play and learning processes.

Key Points: Rich Sutton proposes the OAK architecture as a vision for superintelligence driven entirely by experience, emphasizing that the path to AI lies in improving reinforcement learning algorithms. OAK is designed to be domain-general, experiential (learning only at runtime), and open-ended in sophistication, enabling agents to discover abstractions in complex, 'big worlds'. The 'big world' hypothesis posits that the world is vastly larger and more complex than any agent, necessitating runtime learning and adaptation due to non-stationarity and approximation. A core innovation of OAK is the generation of self-posed sub-problems, inspired by natural play and curiosity, to drive the learning of new features and higher-level concepts. Sutton highlights that while planning with temporally extended behaviors (options) is understood, reliable continual deep learning and robust feature generation remain critical open challenges for the architecture. The architecture aims to provide plausible mechanistic answers to fundamental AI questions, including how high-level knowledge is learned from low-level experience and the origin of concepts.

Context: Rich Sutton, a prominent researcher in reinforcement learning, delivered this talk at the RLC 2025 conference. He has dedicated half a century to developing reinforcement learning algorithms and is known for his 'bitter lesson' and 'reward hypothesis' ideas. This presentation introduces his OAK (Options, Knowledge) architecture, a new vision for AI agent design aimed at achieving superintelligence through experience, contrasting it with approaches that rely heavily on pre-programmed knowledge or non-experiential methods like large language models.

Raw markdown version of this recap