# OpenAI is edging us all... Closer to AGI

Source: https://www.youtube.com/watch?v=rEvEXQvo-F8
Recap page: https://rapidrecap.app/video/rEvEXQvo-F8
Generated: 2025-12-12T23:02:24.414+00:00

---
## Quick Overview

OpenAI's GPT-5.2 launch, featuring a 390x efficiency improvement over GPT-03 on the ARC-AGI-1 benchmark, shifts the AI hype cycle back to OpenAI, even as new benchmarks like ARC-AGI-2 show competitors like Gemini 3 Pro achieving higher scores than GPT-5.2 Pro in abstract reasoning, while Disney's $1 billion investment in OpenAI to integrate characters into Sora raises concerns about IP usage in generative video.

**Key Points:**
- GPT-5.2 achieved a 390x efficiency improvement in one year on the ARC-AGI-1 benchmark, moving from an 88% score on ARC-AGI-1 for an unreleased O3 model to 90.5% for GPT-5.2 Pro.
- On the more challenging ARC-AGI-2 benchmark, GPT-5.2 Pro (High) scored 54.2%, while Google's Gemini 3 Pro scored 43.3% and Claude Opus 4.5 scored 52.0%, indicating competitors are leading in abstract reasoning.
- Marc Benioff praised Gemini 3 for its speed and capability across reasoning, speed, images, and video, suggesting a significant leap over previous models.
- GPT-5.2 demonstrated better long-context recall than GPT-4.1, maintaining near 100% match ratio on MRC-Rv2 with 4 needles, while GPT-4.1 showed a significant drop-off.
- GPT-5.2 Thinking showed a 30-40% reduction in major factual errors compared to GPT-4-0 Thinking when browsing was enabled, performing on par or slightly better than predecessors.
- Disney announced a $1 billion equity investment in OpenAI, allowing users to generate videos featuring over 200 Disney, Marvel, Pixar, and Star Wars characters using Sora and ChatGPT Image.
- The video highlights the intense competition in the AI space, comparing the rapid iteration cycle (OpenAI, Google/Gemini, Anthropic, Grok) and the rising importance of benchmarks that test generalization (ARC-AGI-2) over narrow tasks (ARC-AGI-1).

![Screenshot at 00:07: The video highlights Google's Gemini 3 achieving unexpected dominance on performance-to-cost graphs, particularly on benchmarks like ARC-AGI-2, which tests abstract reasoning.](https://ss.rapidrecap.app/screens/rEvEXQvo-F8/00-00-07.png)

**Context:** This video provides a technical update on the rapidly evolving field of Large Language Models (LLMs) in December 2025, focusing specifically on the recent release of OpenAI's GPT-5.2 and how it stacks up against competitors like Google's Gemini 3 and Anthropic's Claude Opus.

## Detailed Analysis

The video analyzes the recent release of OpenAI's GPT-5.2 amidst fierce competition. OpenAI declared a 'Code Red' due to competitive pressure, evidenced by a 6% decline in ChatGPT traffic, partially attributed to Google's Gemini 3 launch. On the ARC-AGI-1 benchmark, GPT-5.2 Pro showed massive efficiency gains, achieving a 390x improvement over the previous year's O3 model, reaching a 90.5% score. However, on the more difficult ARC-AGI-2 benchmark, which tests generalization, GPT-5.2 Pro scored 54.2%, while Claude Opus 4.5 scored 52.0% and Gemini 3 Pro scored 43.3% (though Gemini 3 Deep Think scored higher than GPT-5.2). Marc Benioff strongly endorsed Gemini 3 for its reasoning, speed, and multimodal capabilities. Furthermore, the video discusses Disney's massive $1 billion investment in OpenAI, which grants access to use over 200 Disney, Marvel, and Star Wars characters in Sora video generation, raising significant intellectual property concerns. Finally, GPT-5.2 shows improved long-context recall over GPT-4.1 and claims 30-40% fewer major factual hallucinations than its predecessors. The video concludes by promoting Railway as a cloud platform offering significant cost savings and fast deployment times, suggesting it is a superior alternative to complex setups like AWS.

### AI Model Releases & Competition

- OpenAI declares 'Code Red' due to traffic decline and competitive pressure from Google's Gemini 3
- AI Hype Cycle shifts back to OpenAI with GPT-5.2 launch
- Competition includes Grok, Anthropic, and Google.

### Benchmark Performance (ARC-AGI-1 & 2)

- GPT-5.2 Pro shows a 390x efficiency improvement on ARC-AGI-1 (90.5% score) compared to O3
- ARC-AGI-2 (Abstract Reasoning) shows GPT-5.2 Pro at 54.2%, with Claude Opus 4.5 at 52.0% and Gemini 3 Pro at 43.3%.

### Model Capabilities

- GPT-5.2 demonstrates superior long-context recall compared to GPT-4.1
- GPT-5.2 Thinking reduces major factual errors by 30-40% compared to GPT-4-0 Thinking in browsing-enabled scenarios.

### Industry Moves & IP Concerns

- Disney invests $1 billion in OpenAI, enabling use of 200+ copyrighted characters (Mickey Mouse, Star Wars) in Sora video generation
- This partnership raises concerns about IP usage in generative AI.

### Developer Productivity & Deployment

- Railway is showcased as a platform offering 50% faster build times and cost savings (up to 65% vs. traditional cloud) for deploying applications and infrastructure.

### Prediction Markets & Insider Trading

- The accuracy of prediction markets like Polymarket in forecasting the GPT-5.2 release led to accusations of insider trading involving Google employees.

![Screenshot at 00:04: A chart illustrating the ChatGPT traffic decline between November 11 and December 1, 2025, showing a 6% decline from the peak.](https://ss.rapidrecap.app/screens/rEvEXQvo-F8/00-00-04.png)
![Screenshot at 00:07: The ARC-AGI-2 Leaderboard comparing various models on cost per task vs. score percentage, showing GPT-5.2 Pro and Gemini 3 Pro clustered in the high-performance region.](https://ss.rapidrecap.app/screens/rEvEXQvo-F8/00-00-07.png)
![Screenshot at 00:09: A tweet from Marc Benioff praising Gemini 3 as an insane leap forward compared to 3 years of using ChatGPT.](https://ss.rapidrecap.app/screens/rEvEXQvo-F8/00-00-09.png)
![Screenshot at 00:36: A tweet from ARC Prize highlighting GPT-5.2 Pro's 90.5% SOTA score on ARC-AGI-1, representing a ~390x efficiency improvement over the O3 model.](https://ss.rapidrecap.app/screens/rEvEXQvo-F8/00-00-36.png)
![Screenshot at 02:55: A graph comparing GPT-5.2 Thinking versus GPT-4.1 Thinking on long-context recall \(Mean match ratio vs. Max input tokens\), showing GPT-5.2 maintaining a much higher score across token lengths.](https://ss.rapidrecap.app/screens/rEvEXQvo-F8/00-02-55.png)
