GPT 5.2 is here.. the biggest update yet!
Quick Overview
OpenAI introduced GPT-5.2, which sets a new state-of-the-art across many benchmarks, notably achieving 90.5% SOTA on ARC-AGI-1 (X-High) at $11.64/task and showing substantial performance gains over GPT-5.1, including a 390X efficiency improvement on ARC-AGI-1 over one year, while also demonstrating significantly reduced hallucination rates (6.2% error for GPT-5.2 vs 8.8% for GPT-5.1 Thinking) and near-perfect long-context reasoning (98% match ratio at 256k tokens on MRCv2).
Key Points: GPT-5.2 Thinking sets a new state-of-the-art on GDPval, achieving 70.9% accuracy, significantly beating GPT-5.1 Thinking's 38.8%. On ARC-AGI-1, GPT-5.2 Pro (X-High) achieved a 90.5% score at $11.64/task, representing a 390X efficiency improvement over an unreleased GPT-o3 (High) model that scored 88% at $4.5k/task a year ago. GPT-5.2 Thinking has a response-level error rate of 6.2% on de-identified ChatGPT queries, compared to 8.8% for GPT-5.1 Thinking, indicating 30% fewer errors. In long-context reasoning (OpenAI MRCv2, 4 needles), GPT-5.2 Thinking maintains near 100% accuracy across input sizes up to 256k tokens, achieving 98% at the maximum tested context. GPT-5.2 Instant is positioned as the capable workhorse for everyday professional tasks like creating spreadsheets, building presentations, and writing code. GPT-5.2 Thinking is designed for deeper work, handling complex tasks like coding, long document summarization, and math/logic steps. GPT-5.2 Pro is the smartest and most trustworthy option for difficult questions where a higher-quality answer justifies the wait, showing stronger performance in complex domains like programming.
Context: OpenAI announced the release of GPT-5.2 across three variants: Instant, Thinking, and Pro, positioning it as their most capable model series yet for professional knowledge work and long-running agents. The announcement details significant performance improvements across various benchmarks compared to GPT-5.1 and other frontier models, focusing on reasoning, long-context understanding, tool use, and safety improvements.