# GPT 5.2 is here.. the biggest update yet!

Source: https://www.youtube.com/watch?v=n8uys5zj28Q
Recap page: https://rapidrecap.app/video/n8uys5zj28Q
Generated: 2025-12-11T21:04:44.951+00:00

---
## Quick Overview

OpenAI introduced GPT-5.2, which sets a new state-of-the-art across many benchmarks, notably achieving 90.5% SOTA on ARC-AGI-1 (X-High) at $11.64/task and showing substantial performance gains over GPT-5.1, including a 390X efficiency improvement on ARC-AGI-1 over one year, while also demonstrating significantly reduced hallucination rates (6.2% error for GPT-5.2 vs 8.8% for GPT-5.1 Thinking) and near-perfect long-context reasoning (98% match ratio at 256k tokens on MRCv2).

**Key Points:**
- GPT-5.2 Thinking sets a new state-of-the-art on GDPval, achieving 70.9% accuracy, significantly beating GPT-5.1 Thinking's 38.8%.
- On ARC-AGI-1, GPT-5.2 Pro (X-High) achieved a 90.5% score at $11.64/task, representing a ~390X efficiency improvement over an unreleased GPT-o3 (High) model that scored 88% at $4.5k/task a year ago.
- GPT-5.2 Thinking has a response-level error rate of 6.2% on de-identified ChatGPT queries, compared to 8.8% for GPT-5.1 Thinking, indicating 30% fewer errors.
- In long-context reasoning (OpenAI MRCv2, 4 needles), GPT-5.2 Thinking maintains near 100% accuracy across input sizes up to 256k tokens, achieving 98% at the maximum tested context.
- GPT-5.2 Instant is positioned as the capable workhorse for everyday professional tasks like creating spreadsheets, building presentations, and writing code.
- GPT-5.2 Thinking is designed for deeper work, handling complex tasks like coding, long document summarization, and math/logic steps.
- GPT-5.2 Pro is the smartest and most trustworthy option for difficult questions where a higher-quality answer justifies the wait, showing stronger performance in complex domains like programming.

![Screenshot at 00:04: A comparison table showing GPT-5.2 Thinking scores \(e.g., 100.0% on AIME 2025\) significantly surpassing GPT-5.1 Thinking across multiple reasoning and math benchmarks.](https://ss.rapidrecap.app/screens/n8uys5zj28Q/00-00-04.png)

**Context:** OpenAI announced the release of GPT-5.2 across three variants: Instant, Thinking, and Pro, positioning it as their most capable model series yet for professional knowledge work and long-running agents. The announcement details significant performance improvements across various benchmarks compared to GPT-5.1 and other frontier models, focusing on reasoning, long-context understanding, tool use, and safety improvements.

## Detailed Analysis

OpenAI introduced GPT-5.2, claiming it is the most advanced frontier model for professional work and long-running agents, rolling out across Instant, Thinking, and Pro tiers. GPT-5.2 Thinking achieves state-of-the-art results on several benchmarks, including GDPval (70.9% vs 38.8% for GPT-5.1 Thinking) and setting new records on ARC-AGI-1. GPT-5.2 Pro (X-High) scored 90.5% on ARC-AGI-1 at $11.64/task, a massive efficiency gain compared to a previous unreleased version from a year ago. The model shows improved factuality, with GPT-5.2 Thinking hallucinating 30% less often than GPT-5.1 Thinking (6.2% error vs 8.8%). Vision capabilities are also strengthened, cutting error rates in half for chart reasoning and software interface understanding. In long-context reasoning (MRCv2), GPT-5.2 maintains near 100% accuracy up to 256k tokens, while GPT-5.1 performance degrades significantly at larger contexts. GPT-5.2 Instant is aimed at everyday work, while Thinking handles deeper, complex tasks, and Pro is the premium option for difficult, high-quality answers.

### Benchmark Performance (Reasoning/Math)

- GPT-5.2 Thinking achieves 100.0% on AIME 2025 (vs 94.0% for GPT-5.1 Thinking); GPT-5.2 Thinking scores 92.4% on GPQA Diamond (vs 88.1% for GPT-5.1 Thinking).

### ARC-AGI-1 Efficiency

- GPT-5.2 Pro (X-High) achieved 90.5% SOTA at $11.64/task, a 390X efficiency improvement over a previous model that scored 88% at $4.5k/task a year prior.

### ARC-AGI-2 Performance

- GPT-5.2 Pro (High) scored 54.2% at $15.72/task, making it SOTA for that benchmark, though GPT-5.2 Pro X-High verification was limited by API timeouts.

### Factuality/Error Rates

- GPT-5.2 Thinking has a 6.2% response-level error rate on de-identified queries, compared to 8.8% for GPT-5.1 Thinking, showing reduced hallucinations.

### Long Context Reasoning (MRCv2)

- GPT-5.2 Thinking maintains near 100% accuracy across context sizes up to 256k tokens, while GPT-5.1 Thinking performance drops significantly (to ~40% match ratio at 256k tokens).

### Model Tiers & Use Cases

- GPT-5.2 Instant is for everyday work (spreadsheets, presentations); GPT-5.2 Thinking is for deeper work (coding, complex documents); GPT-5.2 Pro is for the most difficult questions requiring the highest quality.

### Vision Improvements

- GPT-5.2 Thinking cuts error rates roughly in half on chart reasoning and software interface understanding tasks.

![Screenshot at 00:04: Comparison table showing GPT-5.2 Thinking significantly outperforming GPT-5.1 Thinking across multiple benchmarks like GPQA Diamond and AIME 2025.](https://ss.rapidrecap.app/screens/n8uys5zj28Q/00-00-04.png)
![Screenshot at 00:18: The ARC-AGI-1 Leaderboard graph illustrating GPT-5.2 models achieving high scores at significantly lower costs per task compared to previous models.](https://ss.rapidrecap.app/screens/n8uys5zj28Q/00-00-18.png)
![Screenshot at 05:05: A side-by-side comparison demonstrating GPT-5.2 Thinking's superior ability to create a structured 'Workforce planner' spreadsheet compared to GPT-5.1 Thinking.](https://ss.rapidrecap.app/screens/n8uys5zj28Q/00-05-05.png)
![Screenshot at 06:25: A graph showing GPT-5.2 Thinking maintaining near 100% match ratio in the long-context OpenAI MRCv2 benchmark, even at 256k max input tokens, while GPT-5.1 Thinking degrades.](https://ss.rapidrecap.app/screens/n8uys5zj28Q/00-06-25.png)
![Screenshot at 10:08: A table comparing mental health evaluation scores, showing GPT-5.2 models \(Instant and Thinking\) scoring higher than GPT-5.1 counterparts, indicating improved safety responses.](https://ss.rapidrecap.app/screens/n8uys5zj28Q/00-10-08.png)
