Maxim Lott: AI progress continues, as IQ scores rise linearly
Quick Overview
AI progress, as evidenced by rising IQ scores on reasoning tests, continues its linear trajectory, defying speculation that growth had plateaued, with top models like GPT-4 and Claude 3 Opus consistently outperforming prior versions and even human benchmarks like college students.
Key Points: The consistent stream of AI progress, noted in publications like The New Yorker and WSJ, suggests that the rapid improvement curve in Large Language Models (LLMs) is not slowing down. The speaker references data from Tracking AI showing that the measured AI IQ score has risen consistently over the last 18 months, suggesting progress is linear rather than hitting a plateau. GPT-4 and Claude 3 Opus are cited as leading models that have significantly outperformed previous versions, with Opus scoring 148 on a reasoning test, near the theoretical maximum of 150. The Norwegian Mensa test, which is publicly available, shows GPT-4 scoring 148 (equivalent to an IQ of 148), demonstrating significant advancement over prior models which scored in the 60s. The speaker suggests that the real risk is not AI achieving human-level reasoning, but companies secretly dialing back compute to save money, which would artificially stall progress. The current linear progression suggests that by late 2026 or early 2027, AI models will match or exceed human cognitive ability in reasoning and potentially visual tasks.
Context: The discussion centers on recent benchmarks and performance metrics for cutting-edge Large Language Models (LLMs), particularly OpenAI's GPT series and Anthropic's Claude models. The speakers analyze data tracking AI performance over time to counter the narrative that AI progress has stalled or hit a plateau, using specific IQ test scores to quantify the consistent, linear improvement observed.
Detailed Analysis
The discussion confirms a consistent, ongoing stream of improvement in AI, contrary to recent media narratives suggesting a slowdown or plateau in progress. This progress is quantified using IQ scores derived from standardized tests. The speaker cites data from Tracking AI, noting that the progression has been a steady, linear climb over the past year and a half, rather than exhibiting the sudden jumps previously seen. Specifically, the latest models, like GPT-4 and Claude 3 Opus, show dramatic improvements. For instance, Claude 3 Opus scored 148 on a reasoning test, nearly reaching the theoretical maximum of 150, while earlier models scored much lower (in the 60s) on similar tests. The speaker uses the public Norwegian Mensa test scores as a concrete example, where GPT-4 scored 148, translating to an IQ of 148, significantly better than a typical college student's performance (which would be in the low 120s). The key takeaway is that the progress is real, quantifiable, and driven by several teams pushing the frontier simultaneously, which acts as a check against any single company artificially slowing development. This sustained progress implies that AI systems could match or exceed human cognitive abilities in reasoning and vision within the next two years (by late 2026/early 2027).