# More New AI Models! OpenAI Drops 5.1 Pro and Codex Pro

Source: https://www.youtube.com/watch?v=1hI1Pi11Uzo
Recap page: https://rapidrecap.app/video/1hI1Pi11Uzo
Generated: 2025-11-21T02:02:21.5+00:00

---
## Quick Overview

OpenAI announced GPT-5.1-Codex-Max, a new frontier agentic coding model available in Codex, which is built upon GPT-5.1 and surpasses previous models like GPT-5.1-Codex, showing significant improvements in token efficiency and reasoning capabilities, particularly for long-running tasks via a new native compaction mechanism.

**Key Points:**
- OpenAI introduced GPT-5.1-Codex-Max, a new agentic coding model available in Codex, built on an update to the foundational GPT-5.1 reasoning model.
- GPT-5.1-Codex-Max is faster, more intelligent, and more token-efficient than GPT-5.1-Codex, achieving better performance using 30% fewer thinking tokens on SWE-Bench Verified tasks with 'medium' reasoning effort.
- The model introduces native 'compaction,' a mechanism that prunes and compresses the working session's history to preserve context over long horizons, enabling it to work on tasks for over 24 hours.
- Benchmark results on SWE-Lancer IC SWE showed GPT-5.1-Codex-Max achieving 79.9% accuracy compared to GPT-5.1-Codex's 68.3%, and on Terminal-Bench 2.0, it reached 68.7% accuracy compared to 62.8%.
- The release coincided with the announcement of GPT-5.1 Pro for all Pro users, described as a slow, heavy-weight reasoning model, which is still better than Gemini 3 on some complex tasks but suffers from interface limitations.
- The announcement was preceded by community discussion, including a tweet from Noam Brown noting the ability to work autonomously for more than a day, and a review from Matt Shumer calling 5.1 Pro an 'absolute monster' but noting interface friction.
- The improvements in long-horizon capabilities suggest OpenAI is pushing the frontier toward more general, reliable AI systems, as demonstrated by the METR chart showing tripled long-time horizon capability since early 2023.

![Screenshot at 00:00: The initial screen displays the OpenAI blog post titled "Building more with GPT-5.1-Codex-Max," introducing the new frontier agentic coding model and providing context for the subsequent discussion.](https://ss.rapidrecap.app/screens/1hI1Pi11Uzo/00-00-00.png)

**Context:** The video summarizes recent announcements and community reactions concerning OpenAI's new models, specifically GPT-5.1-Codex-Max and GPT-5.1 Pro, released around November 19, 2025. GPT-5.1-Codex-Max is positioned as a significant upgrade to the Codex ecosystem, focusing on enhanced coding capabilities, especially for long-running, complex tasks, achieved through a novel 'compaction' technique to manage context windows effectively. The context is set against the backdrop of perceived competition, particularly with Google's Gemini 3.

## Detailed Analysis

OpenAI launched GPT-5.1-Codex-Max, an update to the Codex model, emphasizing its improved capabilities for long-horizon coding tasks through a new native compaction mechanism. This feature allows the model to prune and compress its working session history, maintaining context over extended operations (observed working for over 24 hours). The announcement was made on November 19, 2025. Benchmarks, specifically SWE-Lancer IC SWE and Terminal-Bench 2.0, show GPT-5.1-Codex-Max significantly outperforming the previous GPT-5.1-Codex model in accuracy, often achieving superior results with less computational effort (30% fewer thinking tokens at 'medium' reasoning effort). Separately, OpenAI released GPT-5.1 Pro to all ChatGPT Pro users, which was met with mixed reviews; while powerful, some users noted it was slower than previous versions and suffered from interface friction, living only in ChatGPT rather than IDEs. Community figures like Noam Brown and Ethan Mollick highlighted the model's autonomy and significant jump in ability, while tech analyst Gavin Baker's data reinforced the rapid acceleration of AI capabilities in handling long-horizon engineering tasks, showing Codex-Max leading the pack into 2025.

### GPT-5.1-Codex-Max Introduction

- Introducing GPT-5.1-Codex-Max, a new frontier agentic coding model available in Codex
- Built on an update to the foundational GPT-5.1 reasoning model
- Faster, more intelligent, and more token-efficient than GPT-5.1-Codex

### Compaction Mechanism

- Enables completion of tasks previously failing due to context-window limits, such as complex refactors and long-running agent loops
- Prunes and compresses working session history while preserving important context over long horizons
- Model worked on tasks for more than 24 hours in internal evaluations

### Frontier Coding Capabilities & Benchmarks

- Trained on real-world software engineering tasks (PR creation, code review, frontend coding, Q&A)
- Outperforms previous models on many frontier coding evaluations
- SWE-Lancer IC SWE: 79.9% (Codex-Max) vs 68.3% (Codex)
- Terminal-Bench 2.0: 68.7% (Codex-Max) vs 62.8% (Codex)

### GPT-5.1 Pro Review (Matt Shumer)

- Described as a 'fucking monster,' but not all positive
- Biggest weakness is the interface (lives in ChatGPT, not IDE)
- Slower, heavy-weight reasoning model compared to Gemini 3 on some tasks

### General AI Progress (Gavin Baker)

- Gemini 3 shows scaling laws for pretraining are intact
- Blackwell models show a significant performance increase
- GPT-5.1 Pro is a 10-15% jump over GPT-5 Pro for certain use cases, feels like a step toward real colleagues

![Screenshot at 00:00: The OpenAI announcement page for GPT-5.1-Codex-Max, showcasing the title and the 'Get started' button.](https://ss.rapidrecap.app/screens/1hI1Pi11Uzo/00-00-00.png)
![Screenshot at 02:17: A comparison chart showing GPT-5.1-Codex vs GPT-5.1-Codex-Max accuracy on SWE-Lancer IC SWE \(79.9% vs 68.3%\) and Terminal-Bench 2.0 \(68.7% vs 62.8%\).](https://ss.rapidrecap.app/screens/1hI1Pi11Uzo/00-02-17.png)
![Screenshot at 03:03: The section on 'Long-running tasks' explaining how compaction enables GPT-5.1-Codex-Max to handle tasks over 24 hours.](https://ss.rapidrecap.app/screens/1hI1Pi11Uzo/00-03-03.png)
![Screenshot at 04:43: The METR chart illustrating the time-horizon of software engineering tasks completed by different LLMs, with GPT-5.1-Codex-Max showing the largest capability jump by 2025.](https://ss.rapidrecap.app/screens/1hI1Pi11Uzo/00-04-43.png)
![Screenshot at 05:55: A tweet from Alex Volkov summarizing the release, noting the native compaction mechanism and the jump in capability over GPT-5.1.](https://ss.rapidrecap.app/screens/1hI1Pi11Uzo/00-05-55.png)
