# Anthropic's Claude Opus 4.5 in 5 Minutes

Source: https://www.youtube.com/watch?v=TrouQWADTU4
Recap page: https://rapidrecap.app/video/TrouQWADTU4
Generated: 2025-11-24T23:03:19.526+00:00

---
## Quick Overview

Anthropic introduced Claude Opus 4.5, positioning it as the best model for coding, agents, and computer use, demonstrated by its state-of-the-art performance on software engineering benchmarks, including leading Opus 4.5 with 80.9% accuracy on SWE-bench Verified, and achieving 89.3% on Agentic terminal coding.

**Key Points:**
- Claude Opus 4.5 is available now, claimed to be the best model for coding, agents, and computer use.
- Opus 4.5 achieved 80.9% accuracy on the SWE-bench Verified benchmark, outperforming competitors like Sonnet 4.5 (77.2%) and Gemini 1.5 Pro (76.2%).
- The model also scored 89.3% on Agentic terminal coding, significantly ahead of Sonnet 4.5 (50.0%).
- Pricing for the Claude API with Opus 4.5 is $5/$25 per million tokens for input/output.
- New features include an effort parameter on the Claude API to control time/spend tradeoff, and updates to Claude Code, including parallel local/remote sessions.
- Early testers noted Opus 4.5 handles ambiguity and trade-offs without hand-holding, fixing complex multi-system bugs.
- Opus 4.5 scored higher than any human candidate ever on a notoriously difficult performance engineering take-home exam.

![Screenshot at 00:17: A bar chart comparing Opus 4.5 performance against competitors \(Sonnet 4.5, Opus 4.1, Gemini 1.5 Pro, GPT-5.1\) on Software Engineering benchmarks, showing Opus 4.5 leading with 80.9% accuracy on Agentic coding.](https://ss.rapidrecap.app/screens/TrouQWADTU4/00-00-17.png)

**Context:** Anthropic announced the release of its newest large language model, Claude Opus 4.5, on November 24, 2025. The announcement detailed significant performance improvements across various tasks, particularly in coding, agentic workflows, and general computer use, benchmarking it against previous Claude versions and competitors like Gemini and GPT models. The company also highlighted new developer platform features that leverage the model's enhanced capabilities.

## Detailed Analysis

Anthropic launched Claude Opus 4.5, asserting it is the most intelligent, efficient, and best model globally for coding, agents, and computer tasks, while also being better at everyday tasks like deep research and spreadsheet work. Benchmarks show Opus 4.5 achieving state-of-the-art results on real-world software engineering tests; specifically, it scored 80.9% on SWE-bench Verified, surpassing Sonnet 4.5 (77.2%) and Gemini 1.5 Pro (76.2%), and achieving 89.3% on Agentic terminal coding. For API users, the pricing is set at $5 per million input tokens and $25 per million output tokens. Furthermore, Anthropic introduced an 'effort control' parameter in the API, allowing users to balance time spent versus capability maximization. The model also demonstrated significant improvements in managing teams of subagents, boosting performance in deep research evaluations by almost 15 percentage points. Product updates coinciding with the release include the availability of Claude Code on desktop apps, enabling parallel sessions for coding, research, and updates, and an upgrade to Plan Mode. Early testers praised Opus 4.5's ability to handle ambiguity and trade-offs without constant guidance, and it notably scored higher than any human candidate on a difficult performance engineering take-home exam.

### Opus 4.5 Benchmarks

- Opus 4.5 leads SWE-bench Verified (80.9%) and Agentic terminal coding (89.3%)
- Opus 4.5 beats GPT-5.1 on Agentic coding (80.9% vs 76.3%) and Terminal coding (89.3% vs 58.1%)
- Opus 4.5 performs better than all previous versions across measured benchmarks.

### Pricing and Efficiency

- Opus 4.5 API pricing is $5/$25 per million tokens
- At the highest effort level, Opus 4.5 exceeds Sonnet 4.5 performance by 4.3 percentage points while using 48% fewer tokens.

### New Developer Features

- Introduction of an effort control parameter in the Claude API to minimize time/spend or maximize capability
- Claude Code is now available on desktop, supporting parallel local/remote sessions for code/research/updates.

### Agentic Performance

- Context management and memory capabilities boost performance on agentic tasks
- Opus 4.5 effectively manages complex, well-coordinated multi-agent systems, boosting deep research evaluation by almost 15 percentage points.

### Product Updates

- Claude Code gains upgrades to Plan Mode for more precise planning and autonomous execution
- Claude for Chrome and Claude for Excel extensions are now available to all Max users, leveraging Opus 4.5 capabilities.

### First Impressions & Testing

- Early testers found Opus 4.5 handles ambiguity and complex bugs without hand-holding
- Opus 4.5 scored higher than any human candidate on a prescribed 2-hour performance engineering take-home exam.

![Screenshot at 00:09: A bar chart illustrating software engineering accuracy scores for Opus 4.5 compared to other models like Sonnet 4.5 and Gemini 1.5 Pro.](https://ss.rapidrecap.app/screens/TrouQWADTU4/00-00-09.png)
![Screenshot at 00:43: A detailed comparison table showing Opus 4.5 achieving the highest accuracy across multiple software engineering benchmarks like Agentic coding \(80.9%\) and Agentic tool use \(88.9%\).](https://ss.rapidrecap.app/screens/TrouQWADTU4/00-00-43.png)
![Screenshot at 01:25: A bar chart showing Long-term coherence performance on the Vending-Bench, where Opus 4.5 achieved $4,967.06 compared to Sonnet 4.5's $3,849.34.](https://ss.rapidrecap.app/screens/TrouQWADTU4/00-01-25.png)
![Screenshot at 02:01: A demonstration of Opus 4.5 solving a 'Puzzle Room Challenge' showing its programmatic tool calling capability succeeding with a 'SUCCESS!' message.](https://ss.rapidrecap.app/screens/TrouQWADTU4/00-02-01.png)
![Screenshot at 03:32: The Claude Code model selection menu within the desktop app, highlighting Opus 4.5 as the default recommended model for complex work.](https://ss.rapidrecap.app/screens/TrouQWADTU4/00-03-32.png)
