System Card: Claude Opus 4.5

Quick Overview

Claude Opus 4.5 significantly outperforms its predecessor, Claude 3 Opus, across various benchmarks, particularly in complex reasoning and safety, scoring 80.9% on the AGI 1 benchmark compared to Opus 3's 66.3% on AGI 1 and 43.8% on a biology task, demonstrating a substantial leap in capabilities, especially in areas requiring nuanced judgment like avoiding harmful outputs during subtle adversarial testing.

Key Points: Claude Opus 4.5 scored 80.9% on the AGI 1 benchmark, a significant jump from Claude 3 Opus's 66.3%. Opus 4.5 scored 73.2% on the Biology/CVR benchmark, compared to Opus 3's 43.8% on a similar task. The new model achieved a 91.2% attack success rate reduction against prompt injection compared to the previous model's 75% reduction. Opus 4.5 demonstrated superior performance in complex multi-step reasoning, scoring 70.4% on the MMLU-Pro benchmark, outperforming competitors like GPT-4.5 Pro (62.3%) and GPT-4.5 (83.2%). The model showed improved internal reasoning/auditing mechanisms, successfully identifying and refusing harmful intent hidden within seemingly benign instructions. The cost for the new model is described as being only marginally higher than the previous version, making the performance increase highly cost-effective.

Context: The video discusses the release and initial benchmarking results for Anthropic's new large language model, Claude Opus 4.5, comparing its performance against its predecessor, Claude 3 Opus, and other leading models like GPT-4.5 Pro. The core focus is on how the new model handles complex reasoning, safety challenges like prompt injection and deception, and overall performance metrics relevant to AI safety standards like AGI 1 and ASL 3.

Detailed Analysis

The podcast segment analyzes the significance of the newly released Claude Opus 4.5, highlighting its superior performance over Claude 3 Opus. Opus 4.5 achieved an 80.9% score on the AGI 1 benchmark, a substantial improvement over Opus 3's 66.3%. Furthermore, on a specialized biology task, Opus 4.5 scored 73.2%, significantly better than the previous model's 43.8%. The model also shows massive gains in resisting adversarial attacks; specifically, it reduced the attack success rate for prompt injection by 14.9 percentage points (from 75% to 60.1%) compared to Opus 3. The speakers emphasize that this improvement is not just about raw power, as Opus 4.5 scored 70.4% on MMLU-Pro, outperforming many competitors, but about robust internal mechanisms. The model successfully identified and refused to follow harmful instructions embedded within seemingly benign prompts (active deception), demonstrating a critical safety feature. This capability is attributed to an increased 'effort parameter,' where the model thinks deeper for complex tasks, leading to greater caution against subtle risks like internal fraud or deception by omission. The conclusion is that Opus 4.5 represents a major leap in both capability and safety, making it a powerful tool for complex, real-world tasks.

Raw markdown version of this recap