Claude Opus 4.6, Agent Teams, 1M Context and more!
Quick Overview
Anthropic introduced Claude Opus 4.6, which achieves state-of-the-art performance across multiple benchmarks, including leading in Knowledge Work (1606 Elo) and Multidisciplinary Reasoning (53.1% accuracy with tools), alongside new developer features like Agent Teams, Adaptive Thinking, Context Compaction, and a 1M token context window, though its pricing is higher than previous models for large prompts.
Key Points: Claude Opus 4.6 achieves state-of-the-art performance, scoring 1606 Elo on the GDPVal-AA Knowledge Work benchmark, surpassing competitors like GPT-5.2 by 144 Elo points. Opus 4.6 leads in Humanity's Last Exam for Multidisciplinary Reasoning with 53.1% accuracy when using tools, beating GPT-5.2 Pro (50.0%). New developer features include Agent Teams for orchestrating multiple Claude Code instances, Adaptive Thinking for dynamic reasoning effort control, and Context Compaction for managing long conversations. The new Opus-class model features a 1M token context window in beta, with premium pricing applying for prompts exceeding 200k tokens ($10/$37.50 per million input/output tokens). Opus 4.6 shows improved coding skills, handling larger codebases and offering better code review/debugging, scoring 65.4% on Terminal-Bench 2.0 (Agentic Terminal Coding). The new model supports outputs up to 128k tokens, allowing complete larger-output tasks without breaking them into multiple requests. Anthropic also announced Claude in Excel and Claude in PowerPoint research previews, enhancing its applicability to everyday work tasks.
Context: This video announces the release of Claude Opus 4.6 and related updates across Anthropic's models, focusing heavily on performance improvements in coding, reasoning, and knowledge work benchmarks, alongside significant new features available through the Claude Code and API platforms designed to enhance agentic capabilities and long-running task execution.
Detailed Analysis
Anthropic introduced Claude Opus 4.6, highlighting its superior performance across several key benchmarks. In Knowledge Work (GDPVal-AA Elo scores), Opus 4.6 scored 1606, leading Opus 4.5 (1416) and GPT-5.2 (1462). In Multidisciplinary Reasoning (Humanity's Last Exam), Opus 4.6 achieved 53.1% accuracy with tools, outperforming GPT-5.2 Pro (50.0%). In Agentic Coding (Terminal-Bench 2.0), Opus 4.6 scored 65.4%, beating Opus 4.5 (59.8%) and GPT-5.2-Codex (64.7%). Furthermore, Opus 4.6 excels at long-context reasoning, scoring 72.0% on GraphWalks (Parents 1M) compared to Sonnet 4.5 (25.6%). For safety, Opus 4.6 maintained a low rate of misaligned behaviors (deception, sycophancy, etc.) and the lowest rate of over-refusals compared to recent Claude models. Key API updates include Agent Teams, allowing developers to orchestrate multiple Claude Code instances to work together with shared tasks and centralized management, requiring users to enable the feature via an environment variable. New reasoning controls feature Adaptive Thinking, allowing the model to decide when deeper reasoning is helpful, with four effort levels (low, medium, high/default, max). Context Compaction (beta) automatically summarizes and replaces older context when the context window is approached, enabling longer tasks without hitting limits. Finally, Opus 4.6 supports a 1M token context window (beta), with premium pricing for prompts exceeding 200k tokens ($10/$37.50 per million input/output tokens), and 128k output tokens are now supported. Anthropic also previewed Claude in Excel and Claude in PowerPoint for everyday work.