Claude just beat Gemini 3... how?!

Quick Overview

Claude Opus 4.5 outperforms Gemini 3 Pro across several benchmarks, achieving state-of-the-art (SOTA) results on the ARC-AGI-1 test (80.9% agentic coding, 98.2% agentic tool use in Telecom) and showing strong performance in long-term coherence (Vending-Bench 2) and sophisticated agentic capabilities, though Anthropic notes it has not yet reached the AI R&D-4 autonomy threshold.

Key Points: Claude Opus 4.5 achieved SOTA results on ARC-AGI-1, scoring 80.9% on agentic coding (SWE-bench Verified) and 98.2% on agentic tool use (Telecom). The model demonstrated strong long-term coherence on Vending-Bench 2, earning $4,967.06 (mean), second only to Gemini 3 Pro ($4,387.93 mean, but Opus 4.5's score is higher in this specific Vending-Bench 2 leaderboard shown). Opus 4.5 outperformed Gemini 3 Pro on ARC-AGI-1 benchmarks, such as Novel Problem Solving (37.6% vs. 31.1%) and Graduate-level Reasoning (87.0% vs. 91.9% for Gemini 3 Pro, indicating a slight lead for Gemini in that specific area). Anthropic's internal testing confirmed Opus 4.5 scored higher than any human candidate on an internal performance engineering take-home exam within a 2-hour limit. Multi-agent configurations using Opus 4.5 as the orchestrator consistently outperformed single-agent baselines in Search Performance (Internal Benchmark), yielding up to an 87.0% score with Haiku 4.5 subagents. The paper highlights that while Opus 4.5 is highly capable, it has not yet reached the AI R&D-4 autonomy threshold, which requires the ability to fully automate an entry-level, remote-only Researcher role at Anthropic. New features like Claude for Chrome and Claude for Excel are expanding Opus 4.5's utility in real-world tasks, including using spreadsheets and handling long-running tasks.

Context: The video analyzes the recent performance benchmarks and new features released by Anthropic for their Claude Opus 4.5 model, comparing it against competitors like Gemini 3 Pro and previous Claude versions across various technical and agentic evaluations. Key evaluations covered include ARC-AGI, agentic benchmarks (like SWE-bench), and long-term coherence testing (Vending-Bench 2), alongside discussions on AI safety and autonomy thresholds.

Raw markdown version of this recap