Anthropic's Claude Opus 4.5 in 5 Minutes
Quick Overview
Anthropic introduced Claude Opus 4.5, positioning it as the best model for coding, agents, and computer use, demonstrated by its state-of-the-art performance on software engineering benchmarks, including leading Opus 4.5 with 80.9% accuracy on SWE-bench Verified, and achieving 89.3% on Agentic terminal coding.
Key Points: Claude Opus 4.5 is available now, claimed to be the best model for coding, agents, and computer use. Opus 4.5 achieved 80.9% accuracy on the SWE-bench Verified benchmark, outperforming competitors like Sonnet 4.5 (77.2%) and Gemini 1.5 Pro (76.2%). The model also scored 89.3% on Agentic terminal coding, significantly ahead of Sonnet 4.5 (50.0%). Pricing for the Claude API with Opus 4.5 is $5/$25 per million tokens for input/output. New features include an effort parameter on the Claude API to control time/spend tradeoff, and updates to Claude Code, including parallel local/remote sessions. Early testers noted Opus 4.5 handles ambiguity and trade-offs without hand-holding, fixing complex multi-system bugs. Opus 4.5 scored higher than any human candidate ever on a notoriously difficult performance engineering take-home exam.
Context: Anthropic announced the release of its newest large language model, Claude Opus 4.5, on November 24, 2025. The announcement detailed significant performance improvements across various tasks, particularly in coding, agentic workflows, and general computer use, benchmarking it against previous Claude versions and competitors like Gemini and GPT models. The company also highlighted new developer platform features that leverage the model's enhanced capabilities.
Detailed Analysis
Anthropic launched Claude Opus 4.5, asserting it is the most intelligent, efficient, and best model globally for coding, agents, and computer tasks, while also being better at everyday tasks like deep research and spreadsheet work. Benchmarks show Opus 4.5 achieving state-of-the-art results on real-world software engineering tests; specifically, it scored 80.9% on SWE-bench Verified, surpassing Sonnet 4.5 (77.2%) and Gemini 1.5 Pro (76.2%), and achieving 89.3% on Agentic terminal coding. For API users, the pricing is set at $5 per million input tokens and $25 per million output tokens. Furthermore, Anthropic introduced an 'effort control' parameter in the API, allowing users to balance time spent versus capability maximization. The model also demonstrated significant improvements in managing teams of subagents, boosting performance in deep research evaluations by almost 15 percentage points. Product updates coinciding with the release include the availability of Claude Code on desktop apps, enabling parallel sessions for coding, research, and updates, and an upgrade to Plan Mode. Early testers praised Opus 4.5's ability to handle ambiguity and trade-offs without constant guidance, and it notably scored higher than any human candidate on a difficult performance engineering take-home exam.