Adam Marblestone – AI is missing something fundamental about the brain

Quick Overview

The fundamental missing piece in current AI, particularly LLMs, compared to the human brain, likely lies in the highly specific and evolutionarily complex loss functions the brain utilizes, rather than just architecture or simple next-token prediction objectives used in machine learning.

Key Points: The speaker posits that current LLMs, despite massive data, possess only a small fraction of human capabilities, suggesting a fundamental element is missing in AI design. The speaker's hunch is that the field has neglected the role of very specific loss functions, contrasting them with the simple, mathematically simple ones like cross-entropy used in ML. A possibility for the cortex is that it functions as an 'incredibly general prediction engine' capable of 'omnidirectional inference,' predicting any subset of variables from any other subset, unlike LLMs which natively predict only the next token. Steve Byrnes' theory suggests the brain separates into a 'Learning Subsystem' (cortex) and a 'Steering Subsystem' (innate responses/rewards), where the cortex learns to predict the Steering Subsystem's outputs. The genomic evidence suggests far more diverse and specialized cell types exist in the Steering Subsystem compared to the more repeating, general architecture of the Learning Subsystem. LLMs use a 'dumb' form of RL conceptually, lacking value functions, whereas parts of the brain like the striatum and basal ganglia likely perform model-free RL, while the cortex engages in model-based reasoning.

Context: The discussion centers on addressing the 'million-dollar question': why current AI models, like Large Language Models (LLMs), significantly lag behind human capabilities despite being trained on vastly more data. The speaker explores this gap by comparing current deep learning paradigms (architecture, learning algorithms, cost functions) against the known and hypothesized structures of the human brain, drawing heavily on concepts synthesized by neuroscientist Steve Byrnes, who proposes a dual system involving a general Learning Subsystem and an innate Steering Subsystem.

Raw markdown version of this recap