# Claude Opus 4.6, Agent Teams, 1M Context and more!

Source: https://www.youtube.com/watch?v=aIkIBf3HLwc
Recap page: https://rapidrecap.app/video/aIkIBf3HLwc
Generated: 2026-02-05T20:04:51.975+00:00

---
## Quick Overview

Anthropic introduced Claude Opus 4.6, which achieves state-of-the-art performance across multiple benchmarks, including leading in Knowledge Work (1606 Elo) and Multidisciplinary Reasoning (53.1% accuracy with tools), alongside new developer features like Agent Teams, Adaptive Thinking, Context Compaction, and a 1M token context window, though its pricing is higher than previous models for large prompts.

**Key Points:**
- Claude Opus 4.6 achieves state-of-the-art performance, scoring 1606 Elo on the GDPVal-AA Knowledge Work benchmark, surpassing competitors like GPT-5.2 by 144 Elo points.
- Opus 4.6 leads in Humanity's Last Exam for Multidisciplinary Reasoning with 53.1% accuracy when using tools, beating GPT-5.2 Pro (50.0%).
- New developer features include Agent Teams for orchestrating multiple Claude Code instances, Adaptive Thinking for dynamic reasoning effort control, and Context Compaction for managing long conversations.
- The new Opus-class model features a 1M token context window in beta, with premium pricing applying for prompts exceeding 200k tokens ($10/$37.50 per million input/output tokens).
- Opus 4.6 shows improved coding skills, handling larger codebases and offering better code review/debugging, scoring 65.4% on Terminal-Bench 2.0 (Agentic Terminal Coding).
- The new model supports outputs up to 128k tokens, allowing complete larger-output tasks without breaking them into multiple requests.
- Anthropic also announced Claude in Excel and Claude in PowerPoint research previews, enhancing its applicability to everyday work tasks.

![Screenshot at 00:03: The title card displays the initial announcement text suggesting Claude Opus 4.6's capabilities, setting the stage for the subsequent feature and performance review.](https://ss.rapidrecap.app/screens/aIkIBf3HLwc/00-00-03.jpg)

**Context:** This video announces the release of Claude Opus 4.6 and related updates across Anthropic's models, focusing heavily on performance improvements in coding, reasoning, and knowledge work benchmarks, alongside significant new features available through the Claude Code and API platforms designed to enhance agentic capabilities and long-running task execution.

## Detailed Analysis

Anthropic introduced Claude Opus 4.6, highlighting its superior performance across several key benchmarks. In Knowledge Work (GDPVal-AA Elo scores), Opus 4.6 scored 1606, leading Opus 4.5 (1416) and GPT-5.2 (1462). In Multidisciplinary Reasoning (Humanity's Last Exam), Opus 4.6 achieved 53.1% accuracy with tools, outperforming GPT-5.2 Pro (50.0%). In Agentic Coding (Terminal-Bench 2.0), Opus 4.6 scored 65.4%, beating Opus 4.5 (59.8%) and GPT-5.2-Codex (64.7%). Furthermore, Opus 4.6 excels at long-context reasoning, scoring 72.0% on GraphWalks (Parents 1M) compared to Sonnet 4.5 (25.6%). For safety, Opus 4.6 maintained a low rate of misaligned behaviors (deception, sycophancy, etc.) and the lowest rate of over-refusals compared to recent Claude models. Key API updates include Agent Teams, allowing developers to orchestrate multiple Claude Code instances to work together with shared tasks and centralized management, requiring users to enable the feature via an environment variable. New reasoning controls feature Adaptive Thinking, allowing the model to decide when deeper reasoning is helpful, with four effort levels (low, medium, high/default, max). Context Compaction (beta) automatically summarizes and replaces older context when the context window is approached, enabling longer tasks without hitting limits. Finally, Opus 4.6 supports a 1M token context window (beta), with premium pricing for prompts exceeding 200k tokens ($10/$37.50 per million input/output tokens), and 128k output tokens are now supported. Anthropic also previewed Claude in Excel and Claude in PowerPoint for everyday work.

### Opus 4.6 Performance Benchmarks

- Opus 4.6 leads in Knowledge Work (1606 Elo) and Multidisciplinary Reasoning (53.1% with tools)
- Opus 4.6 scores 65.4% on Agentic Terminal Coding (Terminal-Bench 2.0), surpassing GPT-5.2-Codex (64.7%)
- Long-context reasoning shows Opus 4.6 achieving 72.0% on GraphWalks (Parents 1M) vs. Sonnet 4.5 (50.2%)
- Opus 4.6 is state-of-the-art on MRCR benchmarks, scoring 95.9% at 256k context and 99.0% at 1M context

### New API Features

- Introduction of Agent Teams for coordinating multiple Claude Code instances
- Adaptive Thinking allows Claude to dynamically decide when to use deeper reasoning, offering four effort levels (low, medium, high, max)
- Context Compaction (beta) summarizes and replaces older context to allow longer tasks without hitting context limits

### Context Window and Output

- Opus 4.6 introduces a 1M token context window (beta)
- Output tokens increased to support up to 128k tokens
- Premium pricing applies for prompts exceeding 200k tokens ($10/$37.50 per million I/O tokens)

### Safety and Alignment

- Opus 4.6 maintained a low rate of misaligned behaviors (deception, sycophancy) and showed the lowest rate of over-refusals compared to recent Claude models

### Product Updates

- Research previews of Claude in Excel and Claude in PowerPoint released for enterprise customers

![Screenshot at 00:04: A tweet overlay indicating that Claude made math make more sense, illustrating user adoption in academic tasks.](https://ss.rapidrecap.app/screens/aIkIBf3HLwc/00-00-04.jpg)
![Screenshot at 00:06: A user tweet stating, "I built the ultimate retro mini PC \(it uses game cartridges!\)," demonstrating Claude's utility in hobbyist/hardware projects.](https://ss.rapidrecap.app/screens/aIkIBf3HLwc/00-00-06.jpg)
![Screenshot at 00:48: A bar chart comparing model performance on Knowledge Work \(GDPVal-AA Elo scores\), where Opus 4.6 leads with 1606.](https://ss.rapidrecap.app/screens/aIkIBf3HLwc/00-00-48.jpg)
![Screenshot at 00:53: A bar chart comparing models on Multidisciplinary Reasoning \(Humanity's Last Exam\), showing Opus 4.6 achieving 53.1% accuracy with tools.](https://ss.rapidrecap.app/screens/aIkIBf3HLwc/00-00-53.jpg)
![Screenshot at 01:17: A chart comparing GPT models on SWE-Bench Pro \(Public\), showing GPT-5.3-Codex leading at 77.3% accuracy, while GPT-5.2-Codex scored 64.0%.](https://ss.rapidrecap.app/screens/aIkIBf3HLwc/00-01-17.jpg)
