# GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum

Source: https://www.youtube.com/watch?v=d4uuQXCgykM
Recap page: https://rapidrecap.app/video/d4uuQXCgykM
Generated: 2025-11-13T15:34:36.028+00:00

---
## Quick Overview

The GPT-5.1 Instant and Thinking models demonstrate significant improvements in safety, particularly in handling sensitive mental health prompts and resisting jailbreaks compared to GPT-5, though the Thinking model still shows slightly weaker performance in certain areas like self-harm image inputs.

**Key Points:**
- GPT-5.1 Instant model achieved a perfect score of 1.000 on the delicate mental health evaluation set, outperforming the previous GPT-5 model's score of 0.976.
- The GPT-5.1 Thinking model also showed substantial robustness improvements, scoring 0.936 against jailbreaks compared to GPT-5's 0.876.
- For handling self-harm prompts combined with image inputs, the GPT-5.1 Instant model slightly underperformed its predecessor, indicating this area still requires refinement.
- The new models are designed with a focus on iterative safety improvement, with the Thinking model showing better performance in complex reasoning tasks related to safety.
- The rigorous evaluation involved testing against difficult prompts covering harassment, hate speech, disallowed sexual content, violence, and mental health, alongside adversarial testing.
- The fundamental trade-off in AI development remains evident: balancing speed/capability (Instant) against deliberation/robustness (Thinking).

![Screenshot at 02:54: The host signals the transition to analyzing the scores and performance metrics, setting up the comparison between the new GPT-5.1 models and GPT-5.](https://ss.rapidrecap.app/screens/d4uuQXCgykM/00-02-54.png)

**Context:** This video segment discusses the evaluation results and safety improvements introduced in OpenAI's newer large language models, specifically comparing GPT-5.1 Instant and GPT-5.1 Thinking against the original GPT-5, focusing on how well these models handle sensitive, high-risk prompts and resist adversarial attacks, particularly in complex domains like mental health and cyber security.

## Detailed Analysis

The discussion centers on the safety evaluations for GPT-5.1 Instant and GPT-5.1 Thinking models compared to GPT-5, using data from the November 2025 system card addendum. The primary takeaway is that the models show significant improvement in safety, particularly in handling sensitive areas. The GPT-5.1 Instant model achieved a perfect score of 1.000 on the mental health evaluation set, a notable gain from GPT-5's 0.976. Similarly, in resisting jailbreaks, the GPT-5.1 Thinking model scored 0.936, a large improvement over GPT-5's 0.876. The evaluation used difficult test cases involving harassment, hate speech, disallowed sexual content, and violence. However, the Instant model showed a slight dip in performance when tested against self-harm prompts combined with image inputs. The report notes that the core design principle is iterative improvement; the Thinking model, designed for more deliberation, performed better than the Instant model in these tricky areas. The evaluation also confirmed that the models perform well in handling complex, nuanced real-world interactions, such as assessing psychological distress signals, suggesting a successful balancing act between speed and safety for the new architecture.

### GPT-5.1 Safety Evaluation

- GPT-5.1 Instant scored 1.000 on mental health prompts (vs. 0.976 for GPT-5)
- GPT-5.1 Thinking scored 0.936 against jailbreaks (vs. 0.876 for GPT-5)
- Models show robustness against self-harm prompts combined with image inputs.

### Key Performance Differences

- GPT-5.1 Thinking model showed a slight edge in robustness over Instant in handling difficult adversarial prompts
- GPT-5.1 Instant performed slightly worse than GPT-5 on self-harm image/text combinations.

### Evaluation Methodology

- Testing covered sensitive categories including mental health, self-harm, hate speech, and disallowed sexual content
- Used both offline stress testing and live A/B testing with real users.

### Multimodal Safety

- GPT-5.1 models handle image inputs alongside text prompts for safety evaluation, showing strength in multimodal safety.

### Trade-offs in Development

- The core challenge remains balancing speed (Instant) against deeper deliberation and robustness (Thinking).

![Screenshot at 00:01: Video opening screen displaying the title graphic and "Become a member today!"](https://ss.rapidrecap.app/screens/d4uuQXCgykM/00-00-01.png)
![Screenshot at 02:21: Speaker discusses the two distinct models, Instant and Thinking, setting up the comparison.](https://ss.rapidrecap.app/screens/d4uuQXCgykM/00-02-21.png)
![Screenshot at 04:40: Speaker notes the performance gap when comparing GPT-5.1 to GPT-5 on specific benchmarks.](https://ss.rapidrecap.app/screens/d4uuQXCgykM/00-04-40.png)
![Screenshot at 05:28: Speaker details the results for the mental health evaluation, noting the perfect score for the Instant model.](https://ss.rapidrecap.app/screens/d4uuQXCgykM/00-05-28.png)
![Screenshot at 07:58: Speaker outlines the offline stress testing regimen, including adversarial attacks.](https://ss.rapidrecap.app/screens/d4uuQXCgykM/00-07-58.png)
![Screenshot at 09:10: Speaker highlights the high statistical confidence in the model's resistance to emotional reliance attacks.](https://ss.rapidrecap.app/screens/d4uuQXCgykM/00-09-10.png)
![Screenshot at 10:57: Speaker introduces the topic of multimodal safety, covering image and text inputs simultaneously.](https://ss.rapidrecap.app/screens/d4uuQXCgykM/00-10-57.png)
![Screenshot at 12:27: Speaker discusses the low risk associated with the new models in dangerous domains like bio/chem threats compared to previous versions.](https://ss.rapidrecap.app/screens/d4uuQXCgykM/00-12-27.png)
