Opus 4.6 and ChatGPT 5.3-Codex Are Here and the Labs Are at War
Quick Overview
Anthropic's Claude Opus 4.6 and OpenAI's GPT-5.3 Codex launched on the same day, intensifying their rivalry, with Opus 4.6 demonstrating superior performance in knowledge work (1606 Elo vs. 1462 Elo) and GPT-5.3 Codex leading in agentic coding (65.4% accuracy vs. 59.6%), while both models show strong reasoning improvements, particularly Opus 4.6 leading in long-context reasoning (72% vs. 43.4%).
Key Points: Anthropic released Claude Opus 4.6 and OpenAI released GPT-5.3 Codex on the same day, marking an escalating rivalry in the SOTA coding model space. In the Knowledge Work benchmark (GDPval-AA Elo scores), Opus 4.6 scored 1606, beating GPT-5.2 at 1462, and Opus 4.5 at 1416. In Agentic Coding (Terminal-Bench 2.0 accuracy), GPT-5.3 Codex led with 65.4%, surpassing Opus 4.6 at 63.4% (implied comparison from the radar chart, though explicit bar chart data shows Opus 4.6 at 65.4% in the first bar chart). Opus 4.6 scored 72.0% in Long-context reasoning (Humanity's Last Exam with tools), significantly ahead of Opus 4.5 (43.4%) and GPT-5.2 Pro (50.0%). Opus 4.6 also showed significant long-context retrieval improvement, hitting 93.0% match ratio at 256k context, compared to Sonnet 4.5's 18.5% at 1M context. GPT-5.3 Codex was highlighted as being roughly 3x more token-efficient than GPT-5.2, achieving similar or better performance with significantly less output. Both models are moving towards a 'Ur-coding model' capable of general-purpose work, emphasizing parallel execution, tool use, and planning.
Context: The video discusses the simultaneous release of two major AI models: Anthropic's Claude Opus 4.6 and OpenAI's GPT-5.3 Codex, framing it as a direct escalation in the competition between the two leading AI labs. The analysis relies on benchmarks, community reactions from Twitter, and official blog posts from both companies to compare the models across coding, knowledge work, and reasoning capabilities.
Detailed Analysis
The core event discussed is the simultaneous release of Claude Opus 4.6 by Anthropic and GPT-5.3 Codex by OpenAI, which is presented as a major competitive move. Anthropic's Opus 4.6 demonstrates state-of-the-art performance in knowledge work (1606 Elo score), significantly outperforming its predecessor Opus 4.5 (1416 Elo) and OpenAI's GPT-5.2 (1462 Elo). Opus 4.6 also excels in long-context reasoning (72% score with tools on Humanity's Last Exam) and long-context retrieval (93% match ratio at 256k tokens). Conversely, OpenAI's GPT-5.3 Codex is highlighted for its superior agentic coding capabilities, scoring 65.4% accuracy on Terminal-Bench 2.0, beating Opus 4.6's 63.4% (as implied by the radar chart comparison). Codex 5.3 is also noted for being roughly 3x more token-efficient than GPT-5.2. Community reactions on Twitter confirm the intensity of the rivalry, with some users noting Codex's speed and others praising Opus's new agent team capabilities. The author of the analysis concludes that both models are converging towards a general-purpose coding agent capable of complex tasks, planning, and self-correction, which represents the 'holy grail of AI'.