# GPT 5.2 is scary good...

Source: https://www.youtube.com/watch?v=aNYl-O-XxCA
Recap page: https://rapidrecap.app/video/aNYl-O-XxCA
Generated: 2025-12-12T11:03:49.436+00:00

---
## Quick Overview

GPT-5.2 outperforms GPT-5.1 and older models like Grok 4 across knowledge work tasks (GDPVal at 70.9% vs 38.8%), coding (SWE-Bench Verified at 80.0% vs 76.3%), and science (GPQA Diamond at 92.4% vs 88.1%), demonstrating significant, non-linear progress that challenges previous expectations of AI capability growth and economic impact, as seen in the demonstration of the complex "Spherical Life" simulation and the development of a full 3D game using AI.

**Key Points:**
- GPT-5.2 Thinking achieved a 70.9% win/tie rate on the GDPVal benchmark, more than doubling GPT-5.1 Thinking's 38.8% score on knowledge work tasks.
- On SWE-bench Verified (software engineering), GPT-5.2 Thinking scored 80.0%, significantly ahead of GPT-5.1 Thinking's 76.3%.
- GPT-5.2 achieved 92.4% on GPQA Diamond (no tools), surpassing GPT-5.1's 88.1% score, indicating strong reasoning capabilities.
- The video demonstrates GPT-5.2's ability to generate complex, runnable code for a 3D game using Three.js, showcasing its proficiency in multi-step, integrated programming tasks.
- The OECD's projection of AI capabilities, highlighted by Rob Wiblin, showed an exponential curve that predicted an 'inexplicable, permanent decline' past 2026, which is being contradicted by current SOTA performance.
- The host uses the 'Spherical Life' simulation to illustrate concepts like diminishing returns (cost vs. intelligence) and showcases the complexity of tasks AI can now handle.
- The presenter notes that the speed of improvement, such as the 390x cost reduction in ARC-AGI-1 performance in one year, suggests that current benchmarks might become obsolete quickly.

![Screenshot at 01:05: The video transitions to an OpenAI announcement slide detailing GPT-5.2 as the 'most advanced frontier model for professional work and long-running agents,' setting the context for the subsequent benchmark analysis.](https://ss.rapidrecap.app/screens/aNYl-O-XxCA/00-01-05.png)

**Context:** The video analyzes the release of OpenAI's GPT-5.2, focusing heavily on benchmark results from the GDPVal and ARC-AGI-1 leaderboards to assess its real-world professional capabilities compared to previous models like GPT-5.1 and competitors like Grok. The host uses these metrics, along with demonstrations of complex coding (generating a 3D game) and running simulations, to argue that AI progress is accelerating faster than some external forecasts, particularly those from economists, had predicted. The discussion centers on how these models are moving beyond simple tasks into complex, multi-step professional work.

## Detailed Analysis

The presenter reviews the capabilities of GPT-5.2, contrasting its performance against GPT-5.1 and other models using data from the GDPVal and ARC-AGI-1 benchmarks. On GDPVal (knowledge work tasks), GPT-5.2 achieved a 70.9% win/tie rate, more than doubling GPT-5.1's 38.8%. In coding tasks (SWE-Bench Verified), GPT-5.2 scored 80.0% compared to 76.3% for GPT-5.1. In science reasoning (GPQA Diamond, no tools), GPT-5.2 hit 92.4%, beating GPT-5.1's 88.1%. The host uses a clip from the movie 'The Princess Bride' to illustrate the potential misunderstanding of terms like 'knowing' versus 'accurately predicting' when evaluating AI. The video then shifts to a demonstration where GPT-5.2 quickly generates and iterates upon a complex 3D game using Three.js, showing its capability in multi-file projects, visual adjustments (like reducing bloom effects), and handling complex instructions over multiple turns. The host also discusses the alarming speed of progress, referencing a tweet showing a 390x cost reduction in one year for ARC-AGI-1 performance, suggesting that previous projections of AI capability growth are already outdated and that the industry is entering a phase of rapid, potentially disruptive acceleration.

### GPT-5.2 Performance on GDPVal

- GPT-5.2 Thinking scored 70.9% wins or ties, GPT-5.1 Thinking scored 38.8% (GPT-5)
- GPT-5.2 significantly outperforms GPT-5.1 in knowledge work tasks.

### Software Engineering Benchmarks

- SWE-Bench Pro public accuracy is 55.6% for GPT-5.2 Thinking vs 50.8% for GPT-5.1 Thinking; SWE-Bench Verified is 80.0% vs 76.3%.

### Science Reasoning

- GPQA Diamond (no tools) accuracy is 92.4% for GPT-5.2 vs 88.1% for GPT-5.1.

### Code Generation Demo

- GPT-5.2 successfully created a full 3D game using Three.js based on an initial prompt, demonstrating multi-step code generation and iterative refinement (e.g., adjusting lighting, speed, and visual effects).

### Critique of External Forecasts

- The OECD's projection showing a massive decline in AI capability post-2026 is shown to be contradicted by current exponential progress.

### ARC-AGI-1 Leaderboard Analysis

- GPT-5.2 Pro (X-High) achieved 90.5% score at $11.64/task, a 390x cost reduction from cO3 (High) which scored 88% at $4.5k/task a year prior.

![Screenshot at 00:00: Demonstration of the 'Spherical Life' Conway's Game of Life simulation running on a cube-sphere grid.](https://ss.rapidrecap.app/screens/aNYl-O-XxCA/00-00-00.png)
![Screenshot at 01:04: OpenAI's announcement slide for GPT-5.2, highlighting its focus on professional work and long-running agents.](https://ss.rapidrecap.app/screens/aNYl-O-XxCA/00-01-04.png)
![Screenshot at 01:19: ChatGPT interface showing GPT-5.2 Pro reasoning for 55 minutes and 13 seconds to generate a complex 3D game.](https://ss.rapidrecap.app/screens/aNYl-O-XxCA/00-01-19.png)
![Screenshot at 03:09: Tweet from Noam Brown highlighting GDPVal results where GPT-5.2 Pro \(74.1%\) beats the industry expert parity line \(~50%\).](https://ss.rapidrecap.app/screens/aNYl-O-XxCA/00-03-09.png)
![Screenshot at 10:16: Bar chart comparing GPT-5.2 and GPT-5.1 performance across various benchmarks, showing GPT-5.2 leading in all listed categories.](https://ss.rapidrecap.app/screens/aNYl-O-XxCA/00-10-16.png)
