# Opus 4.6 and ChatGPT 5.3-Codex Are Here and the Labs Are at War

Source: https://www.youtube.com/watch?v=JqpI65aVJ30
Recap page: https://rapidrecap.app/video/JqpI65aVJ30
Generated: 2026-02-06T20:33:12.737+00:00

---
## Quick Overview

Anthropic's Claude Opus 4.6 and OpenAI's GPT-5.3 Codex launched on the same day, intensifying their rivalry, with Opus 4.6 demonstrating superior performance in knowledge work (1606 Elo vs. 1462 Elo) and GPT-5.3 Codex leading in agentic coding (65.4% accuracy vs. 59.6%), while both models show strong reasoning improvements, particularly Opus 4.6 leading in long-context reasoning (72% vs. 43.4%).

**Key Points:**
- Anthropic released Claude Opus 4.6 and OpenAI released GPT-5.3 Codex on the same day, marking an escalating rivalry in the SOTA coding model space.
- In the Knowledge Work benchmark (GDPval-AA Elo scores), Opus 4.6 scored 1606, beating GPT-5.2 at 1462, and Opus 4.5 at 1416.
- In Agentic Coding (Terminal-Bench 2.0 accuracy), GPT-5.3 Codex led with 65.4%, surpassing Opus 4.6 at 63.4% (implied comparison from the radar chart, though explicit bar chart data shows Opus 4.6 at 65.4% in the first bar chart).
- Opus 4.6 scored 72.0% in Long-context reasoning (Humanity's Last Exam with tools), significantly ahead of Opus 4.5 (43.4%) and GPT-5.2 Pro (50.0%).
- Opus 4.6 also showed significant long-context retrieval improvement, hitting 93.0% match ratio at 256k context, compared to Sonnet 4.5's 18.5% at 1M context.
- GPT-5.3 Codex was highlighted as being roughly 3x more token-efficient than GPT-5.2, achieving similar or better performance with significantly less output.
- Both models are moving towards a 'Ur-coding model' capable of general-purpose work, emphasizing parallel execution, tool use, and planning.

![Screenshot at 00:11: The initial comparison chart showing Claude Opus 4.6 \(1606 Elo\) leading Opus 4.5 \(1416 Elo\) and GPT-5.2 \(1462 Elo\) in the 'Knowledge work' benchmark.](https://ss.rapidrecap.app/screens/JqpI65aVJ30/00-00-11.jpg)

**Context:** The video discusses the simultaneous release of two major AI models: Anthropic's Claude Opus 4.6 and OpenAI's GPT-5.3 Codex, framing it as a direct escalation in the competition between the two leading AI labs. The analysis relies on benchmarks, community reactions from Twitter, and official blog posts from both companies to compare the models across coding, knowledge work, and reasoning capabilities.

## Detailed Analysis

The core event discussed is the simultaneous release of Claude Opus 4.6 by Anthropic and GPT-5.3 Codex by OpenAI, which is presented as a major competitive move. Anthropic's Opus 4.6 demonstrates state-of-the-art performance in knowledge work (1606 Elo score), significantly outperforming its predecessor Opus 4.5 (1416 Elo) and OpenAI's GPT-5.2 (1462 Elo). Opus 4.6 also excels in long-context reasoning (72% score with tools on Humanity's Last Exam) and long-context retrieval (93% match ratio at 256k tokens). Conversely, OpenAI's GPT-5.3 Codex is highlighted for its superior agentic coding capabilities, scoring 65.4% accuracy on Terminal-Bench 2.0, beating Opus 4.6's 63.4% (as implied by the radar chart comparison). Codex 5.3 is also noted for being roughly 3x more token-efficient than GPT-5.2. Community reactions on Twitter confirm the intensity of the rivalry, with some users noting Codex's speed and others praising Opus's new agent team capabilities. The author of the analysis concludes that both models are converging towards a general-purpose coding agent capable of complex tasks, planning, and self-correction, which represents the 'holy grail of AI'.

### Release Overview

- Anthropic released Claude Opus 4.6 and OpenAI released GPT-5.3 Codex on the same day, escalating rivalry
- Both labs are aiming for a general-purpose coding agent capable of complex, autonomous workflows.

### Knowledge Work Benchmarks

- Opus 4.6 achieved 1606 Elo in GDPval-AA, leading GPT-5.2 (1462 Elo) and Opus 4.5 (1416 Elo).

### Long Context Performance

- Opus 4.6 achieved 72.0% in long-context reasoning (Graphwalks with tools) and 93.0% in long-context retrieval at 256k context.

### Coding Performance

- GPT-5.3 Codex led in Agentic Coding (Terminal-Bench 2.0) with 65.4% accuracy, compared to Opus 4.6's 63.4% (implied). Codex 5.3 is also noted as being 3x more token-efficient than GPT-5.2.

### Anthropic's Agent Teams

- Opus 4.6 introduced agent teams (e.g., for building a C compiler autonomously) which was praised by users for enabling parallel work and complex task execution.

### OpenAI's Agent Capabilities

- GPT-5.3 Codex focuses on being an agent that can write, review, and deploy code autonomously, achieving high performance benchmarks like 77.3% on Terminal-Bench 2.0.

### Conclusion

- Both models show a convergence toward creating powerful, general-purpose coding agents that handle complex tasks like research, debugging, and deployment without constant human intervention.

![Screenshot at 00:11: The 'Knowledge work' bar chart comparing Opus 4.6 \(1606\) against GPT-5.2 \(1462\) and Opus 4.5 \(1416\).](https://ss.rapidrecap.app/screens/JqpI65aVJ30/00-00-11.jpg)
![Screenshot at 01:52: The 'Agentic coding' bar chart showing GPT-5.3 Codex leading \(65.4%\) over Opus 4.6 \(63.4% implied/65.4% as first bar\) on Terminal-Bench 2.0.](https://ss.rapidrecap.app/screens/JqpI65aVJ30/00-01-52.jpg)
![Screenshot at 04:08: The 'Introducing Claude Opus 4.6' announcement image collage, featuring various use cases like code review and external tools.](https://ss.rapidrecap.app/screens/JqpI65aVJ30/00-04-08.jpg)
![Screenshot at 04:41: A comparison chart showing GPT-5.3 Codex leading GPT-5.2 Codex in accuracy vs. output tokens on the SWE-Bench Pro benchmark.](https://ss.rapidrecap.app/screens/JqpI65aVJ30/00-04-41.jpg)
![Screenshot at 11:17: The radar chart comparing Claude Opus 4.6 and GPT-5.3 Codex capabilities, showing Opus leading in Long-Context Retrieval and GPT-5.3 Codex leading in Coding and Max Output.](https://ss.rapidrecap.app/screens/JqpI65aVJ30/00-11-17.jpg)
