Gemini 3 just got *scary* good

Quick Overview

Gemini 3 Pro is declared the best-performing AI model across multiple benchmarks, significantly outperforming competitors like Claude Sonnet 4.5, Grok 4, GPT-5.1, and Gemini 2.5 Pro in areas like reasoning, coding, and multimodal tasks, particularly excelling in the Vending-Bench 2 simulation where it achieved over $5,200 in net worth starting from $500, and in the ARC-AGI-2 benchmark with a 31.1% score.

Key Points: Gemini 3 Pro is the leading model in the Vending-Bench 2 simulation, ending with a net worth exceeding $5,200 after starting with $500, significantly outperforming the runner-up Claude Sonnet 4.5 ($3,838.74). Gemini 3 Deep Think achieved a 41% score on the Humanity's Last Exam benchmark, surpassing Gemini 3 Pro's 37.5% and all other tested models. On the ARC-AGI-2 benchmark, Gemini 3 Deep Think scored 45.9%, while Gemini 3 Pro scored 31.1%, demonstrating significant capability in symbolic interpretation. Gemini 3 Pro achieved 100% accuracy on the AIME 2025 mathematics test when using code execution, matching Claude Sonnet 4.5's performance but outperforming GPT-4's 94.0%. Gemini 3 Pro is rated the best model for long-term coherence, evidenced by its leading performance in the Vending-Bench 2 test over a year-long simulation. The model exhibits strong qualitative traits, acting as a persistent negotiator by consistently searching for reasonable offers from wholesale suppliers. The new Gemini 3 models, including Deep Think, are shown to be highly capable across a broad range of benchmarks, including coding, math, and reasoning tasks.

Context: This video details the release and initial performance benchmarks of Google's Gemini 3 models, specifically focusing on Gemini 3 Pro and the experimental Gemini 3 Deep Think mode. The content compares these new models against existing frontier models like Claude Sonnet 4.5, Grok 4, and GPT-5.1 across various evaluation suites, including simulated business operations (Vending-Bench 2), reasoning tests (Humanity's Last Exam, ARC-AGI-2), and mathematical challenges (AIME 2025), emphasizing gains in complex reasoning and long-term coherence.

Raw markdown version of this recap