GPT 5.2: OpenAI Strikes Back
Quick Overview
GPT-5.2 Thinking sets a new state-of-the-art across multiple benchmarks, notably outperforming Gemini 1.5 Pro on GPAQ Diamond (92.4% vs 91.9%) and achieving 100% on AIME 2025 (No tools), while showing significant gains in long-context reasoning (near 100% accuracy on 256k tokens in the MRCRv2 test) and coding accuracy on SWE-Bench Pro (55.6% vs 53% for GPT-5.1 Codex-max).
Key Points: GPT-5.2 Thinking achieves a 100.0% score on the AIME 2025 (No tools) benchmark, surpassing Gemini 3 Pro's 95.0%. On GPQA Diamond (No tools), GPT-5.2 Thinking scores 92.4%, slightly ahead of Gemini 3 Pro's 91.9%. GPT-5.2 Thinking achieves 80.0% on SWE-bench Verified, compared to Gemini 3 Pro's 76.2%. In long-context reasoning (MRCRv2, 4 needles), GPT-5.2 Thinking achieves near 100% accuracy at 256k tokens, while GPT-5.1 Thinking drops significantly. On MLE-Bench-30 (no browsing), GPT-5.2 achieves 55% pass@1, beating GPT-5.1 Codex-max's 53%. OpenAI is pricing GPT-5.2 models competitively, with GPT-5.2 Thinking priced at $1.75/1M input tokens and $14/1M output tokens, which is still cheaper than the Opus tier despite higher quality. The developers are focusing on building superintelligence within the next ten years, as stated by Sam Altman (15:09).
Context: This video discusses the announcement and initial performance metrics of OpenAI's new frontier model, GPT-5.2, contrasting its capabilities against competitors like Gemini 3 Pro, Claude Opus, and previous GPT versions. The analysis covers performance across academic benchmarks (like GPQA, AIME, ARC-AGI-1), real-world software engineering tasks (SWE-Bench), long-context understanding (MRCRv2), and pricing/availability.
Detailed Analysis
The video introduces GPT-5.2 as OpenAI's latest frontier model, highlighting its record-breaking results across numerous benchmarks. On the GPT-4 vs. GPT-5.2 Thinking comparison table (08:08), GPT-5.2 excels in many areas, including scoring 100% on AIME 2025 (No tools) compared to Gemini 3 Pro's 95.0%, and achieving 92.4% on GPQA Diamond (No tools) against Gemini 3 Pro's 91.9%. The model also shows significant improvement in coding tasks, scoring 80.0% on SWE-bench Verified versus 76.2% for Gemini 3 Pro (08:00). A key advancement is in long-context reasoning, where GPT-5.2 Thinking maintains near 100% accuracy even with 256k tokens on the MRCRv2 test, significantly outperforming GPT-5.1 Thinking (13:09). Furthermore, on the MLE-Bench-30 (no browsing) chart (14:01), GPT-5.2 achieves 55% pass@1, beating GPT-5.1 Codex-max's 53%. The speaker notes that while some benchmarks show close competition, GPT-5.2 appears superior in complex reasoning (15:51). Regarding pricing, GPT-5.2 Thinking is priced at $1.75/1M input tokens and $14/1M output tokens, which is noted as being cheaper than Claude Opus 4.5 for similar quality (15:34). The video concludes with a quote from Sam Altman expressing high certainty that superintelligence will be built within ten years (15:09).