Claude Opus 4.6 is INSANE! What's New?
Quick Overview
Claude Opus 4.6 significantly outperforms previous models and competitors like Gemini 3 Pro and GPT-5.2 across knowledge work (1606 Elo score), long-context retrieval (93.0% on 256k context), and software failure diagnosis (34.9% accuracy). New developer features include Adaptive Thinking, context compaction, 1M token context (beta), 128k output tokens, and Claude Code updates like Agent Teams and integrations with Excel and PowerPoint.
Key Points: Opus 4.6 achieved a knowledge work Elo score of 1606, surpassing Opus 4.5 (1416) and GPT-5.2 (1462). In long-context retrieval (MRCR v2, 8-needle), Opus 4.6 achieved 93.0% accuracy with 256k context and 76.0% with 1M context. The model excels at software failure diagnosis (OpenRCA) with 34.9% accuracy, significantly better than Opus 4.5 (26.9%). New API features for developers include Adaptive Thinking for dynamic reasoning depth, four effort levels (low, medium, high, max), Context Compaction, and 1M token context (beta). Opus 4.6 supports outputs up to 128k tokens and new capabilities in Claude Code, such as Agent Teams for coordinating multiple agents. Anthropic also released research previews for Claude in Excel (handling long-running tasks) and Claude in PowerPoint (building/editing slides). The pricing for Claude Opus 4.6 remains $5/$25 per million input/output tokens.
Context: This video announces the release of Claude Opus 4.6, Anthropic's latest and most capable model, detailing its performance improvements over previous Claude versions and competitors across several benchmarks. The presenter walks through key evaluation charts, highlights significant new features available via the API, and showcases new product integrations designed to enhance agentic and knowledge work capabilities.
Detailed Analysis
The video introduces Claude Opus 4.6, highlighting its superior performance across multiple benchmarks. In Knowledge Work (GDPVal-AA Elo scores), Opus 4.6 scored 1606, significantly beating Opus 4.5 (1416), Sonnet 4.5 (1277), Gemini 3 Pro (1195), and GPT-5.2 (1462). Agentic tasks also showed improvement; for example, Agentic coding (SWE-bench Verified) reached 80.8%, slightly behind Opus 4.5's 80.9% but ahead of others. In long-context retrieval (MRCR v2, 8-needle), Opus 4.6 achieved 93.0% accuracy at 256k context and 76.0% at 1M context, demonstrating strong resistance to context rot. Furthermore, Opus 4.6 achieved 34.9% accuracy in Software Failure Diagnosis (OpenRCA), outperforming Opus 4.5 (26.9%). New API features include Adaptive Thinking, which lets the model decide on deeper reasoning based on task complexity, and four Effort levels (low, medium, high/default, max) to control thinking depth, which can be dialed down from high to medium to prevent overthinking on simple tasks. Other updates include Context Compaction for summarizing older context, 1M token context (beta), and 128k output tokens. Claude Code now supports Agent Teams for coordinating parallel work. The model also shows substantial upgrades in office tools, with research previews for Claude in Excel and Claude in PowerPoint, capable of building slides and performing complex, long-running analyses automatically.