# Update to GPT-5 System Card: GPT-5.2

Source: https://www.youtube.com/watch?v=J_BrBMTo-do
Recap page: https://rapidrecap.app/video/J_BrBMTo-do
Generated: 2025-12-18T23:34:35.729+00:00

---
## Quick Overview

The release of GPT-5.2 significantly advances AI capabilities, particularly in complex reasoning and safety, as evidenced by its 93.2% score on the GPQA benchmark and a substantial reduction in harmful outputs compared to GPT-5.1, although the model still exhibits some difficulty with abstract reasoning and context adherence in specific tests.

**Key Points:**
- GPT-5.2 achieves a score of 93.2% on the GPQA benchmark, surpassing the previous model, GPT-5.1, which scored 52.9%.
- The new model demonstrated a massive reduction in harmful outputs (suicide/self-harm/mental health distress) down to 0.1% error rate from 17.6% in GPT-5.1.
- GPT-5.2 achieved a 100% score on the MMLU benchmark for high-stakes professional tasks, specifically outperforming its predecessor in areas like finance and legal analysis.
- The model significantly improved in abstract reasoning, scoring 88.8% on the ARC-AGI benchmark, compared to GPT-5.1's lower score.
- However, GPT-5.2 struggled with spatial reasoning, failing to accurately represent a computer motherboard image and incorrectly formatting some outputs when forced to provide an integer.
- The ability to automate end-to-end cyber operations and perform complex, long-context processing (like analyzing 200,000-page filings) is a major strength of the new architecture.
- The discussion concludes that GPT-5.2 represents a major step forward, particularly in safety and complex reasoning, despite minor remaining limitations in certain specialized reasoning tasks.

![Screenshot at 00:06: The screen displays the podcast branding with the call to action "BECOME A MEMBER TODAY!" overlaid on a radar-like graphic, signifying the discussion of a major information drop.](https://ss.rapidrecap.app/screens/J_BrBMTo-do/00-00-06.png)

**Context:** This AI podcast episode discusses the official documentation drop for the new GPT-5.2 frontier model, positioning it not just as an incremental upgrade but as the most advanced series yet. The conversation focuses on comparing GPT-5.2's performance against its predecessor, GPT-5.1, across various benchmarks, specifically highlighting improvements in factual accuracy, complex reasoning, and safety guardrails, while also noting areas where the new model still falls short.

## Detailed Analysis

The speakers are analyzing the release information for the GPT-5.2 model, noting that it is positioned as the most advanced model series yet, not just an incremental update. Key data points show a massive leap in performance: GPT-5.2 scored 93.2% on the GPQA benchmark, compared to GPT-5.1's 52.9%. Furthermore, safety metrics showed significant improvement; harmful outputs related to suicide or self-harm dropped from 17.6% in GPT-5.1 to a negligible 0.1% in GPT-5.2. The model also scored a perfect 100% on the MMLU benchmark for high-stakes professional tasks like finance and legal work, significantly outperforming GPT-5.1. The improvements extend to complex reasoning, with GPT-5.2 scoring 88.8% on the ARC-AGI abstract reasoning benchmark. However, limitations remain; GPT-5.2 failed spatial reasoning tests (e.g., correctly identifying motherboard components from a low-quality image) and struggled to adhere to strict output formats when asked for an integer. The discussion emphasizes that the model's strength lies in its ability to handle complex, long-context tasks like analyzing massive documents and automating end-to-end cyber operations, leading to a profound shift in engineering and business viability.

### GPT-5.2 Performance Metrics

- 93.2% GPQA score (vs 52.9% for GPT-5.1)
- 100% MMLU score for high-stakes professional tasks
- 88.8% ARC-AGI score for abstract reasoning

### Safety Improvements

- Harmful output rate (suicide/self-harm) dropped from 17.6% (GPT-5.1) to 0.1% (GPT-5.2)
- Low hallucination rate (under 1%) on factual checks

### Key Capabilities

- Automation of end-to-end cyber operations
- Handling long-context tasks like 200,000-page filing analysis
- Tool calling for complex multi-step actions

### Identified Limitations

- Struggles with spatial reasoning (e.g., image output)
- Failed to adhere to strict integer output format when instructed

### Comparative Analysis

- Outperforms GPT-5.1 by 11x speed on certain tasks and significantly reduces the error rate on safety benchmarks.

![Screenshot at 00:06: The screen displays the podcast branding with the call to action "BECOME A MEMBER TODAY!" overlaid on a radar-like graphic, signifying the discussion of a major information drop.](https://ss.rapidrecap.app/screens/J_BrBMTo-do/00-00-06.png)
![Screenshot at 02:24: A comparison point showing that GPT-5.2 reliably performs tasks requiring complex sequencing and tool coordination, unlike its predecessor.](https://ss.rapidrecap.app/screens/J_BrBMTo-do/00-02-24.png)
![Screenshot at 04:40: A screen showing the performance difference, contrasting GPT-5.2's 82.2% score against GPT-5.1's 40.3% on the MMLU benchmark for math.](https://ss.rapidrecap.app/screens/J_BrBMTo-do/00-04-40.png)
![Screenshot at 07:28: A graphic illustrating the critical safety floor where GPT-5.2 maintains an 88.8% score on the ARC-AGI benchmark, indicating strong abstract reasoning.](https://ss.rapidrecap.app/screens/J_BrBMTo-do/00-07-28.png)
![Screenshot at 11:21: A visual representation of the waveform/audio activity during the discussion of the model's performance improvements.](https://ss.rapidrecap.app/screens/J_BrBMTo-do/00-11-21.png)
