GPT 5.2 is scary good...

Quick Overview

GPT-5.2 outperforms GPT-5.1 and older models like Grok 4 across knowledge work tasks (GDPVal at 70.9% vs 38.8%), coding (SWE-Bench Verified at 80.0% vs 76.3%), and science (GPQA Diamond at 92.4% vs 88.1%), demonstrating significant, non-linear progress that challenges previous expectations of AI capability growth and economic impact, as seen in the demonstration of the complex "Spherical Life" simulation and the development of a full 3D game using AI.

Key Points: GPT-5.2 Thinking achieved a 70.9% win/tie rate on the GDPVal benchmark, more than doubling GPT-5.1 Thinking's 38.8% score on knowledge work tasks. On SWE-bench Verified (software engineering), GPT-5.2 Thinking scored 80.0%, significantly ahead of GPT-5.1 Thinking's 76.3%. GPT-5.2 achieved 92.4% on GPQA Diamond (no tools), surpassing GPT-5.1's 88.1% score, indicating strong reasoning capabilities. The video demonstrates GPT-5.2's ability to generate complex, runnable code for a 3D game using Three.js, showcasing its proficiency in multi-step, integrated programming tasks. The OECD's projection of AI capabilities, highlighted by Rob Wiblin, showed an exponential curve that predicted an 'inexplicable, permanent decline' past 2026, which is being contradicted by current SOTA performance. The host uses the 'Spherical Life' simulation to illustrate concepts like diminishing returns (cost vs. intelligence) and showcases the complexity of tasks AI can now handle. The presenter notes that the speed of improvement, such as the 390x cost reduction in ARC-AGI-1 performance in one year, suggests that current benchmarks might become obsolete quickly.

Context: The video analyzes the release of OpenAI's GPT-5.2, focusing heavily on benchmark results from the GDPVal and ARC-AGI-1 leaderboards to assess its real-world professional capabilities compared to previous models like GPT-5.1 and competitors like Grok. The host uses these metrics, along with demonstrations of complex coding (generating a 3D game) and running simulations, to argue that AI progress is accelerating faster than some external forecasts, particularly those from economists, had predicted. The discussion centers on how these models are moving beyond simple tasks into complex, multi-step professional work.

Raw markdown version of this recap