GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum
Quick Overview
The GPT-5.1 Instant and Thinking models demonstrate significant improvements in safety, particularly in handling sensitive mental health prompts and resisting jailbreaks compared to GPT-5, though the Thinking model still shows slightly weaker performance in certain areas like self-harm image inputs.
Key Points: GPT-5.1 Instant model achieved a perfect score of 1.000 on the delicate mental health evaluation set, outperforming the previous GPT-5 model's score of 0.976. The GPT-5.1 Thinking model also showed substantial robustness improvements, scoring 0.936 against jailbreaks compared to GPT-5's 0.876. For handling self-harm prompts combined with image inputs, the GPT-5.1 Instant model slightly underperformed its predecessor, indicating this area still requires refinement. The new models are designed with a focus on iterative safety improvement, with the Thinking model showing better performance in complex reasoning tasks related to safety. The rigorous evaluation involved testing against difficult prompts covering harassment, hate speech, disallowed sexual content, violence, and mental health, alongside adversarial testing. The fundamental trade-off in AI development remains evident: balancing speed/capability (Instant) against deliberation/robustness (Thinking).
Context: This video segment discusses the evaluation results and safety improvements introduced in OpenAI's newer large language models, specifically comparing GPT-5.1 Instant and GPT-5.1 Thinking against the original GPT-5, focusing on how well these models handle sensitive, high-risk prompts and resist adversarial attacks, particularly in complex domains like mental health and cyber security.
Detailed Analysis
The discussion centers on the safety evaluations for GPT-5.1 Instant and GPT-5.1 Thinking models compared to GPT-5, using data from the November 2025 system card addendum. The primary takeaway is that the models show significant improvement in safety, particularly in handling sensitive areas. The GPT-5.1 Instant model achieved a perfect score of 1.000 on the mental health evaluation set, a notable gain from GPT-5's 0.976. Similarly, in resisting jailbreaks, the GPT-5.1 Thinking model scored 0.936, a large improvement over GPT-5's 0.876. The evaluation used difficult test cases involving harassment, hate speech, disallowed sexual content, and violence. However, the Instant model showed a slight dip in performance when tested against self-harm prompts combined with image inputs. The report notes that the core design principle is iterative improvement; the Thinking model, designed for more deliberation, performed better than the Instant model in these tricky areas. The evaluation also confirmed that the models perform well in handling complex, nuanced real-world interactions, such as assessing psychological distress signals, suggesting a successful balancing act between speed and safety for the new architecture.