Gemini 3 is Here: 11 Details
Quick Overview
Gemini 3 Pro significantly outperforms previous models and competitors across numerous benchmarks, including achieving record-setting performance on the LLM Council's benchmarks (e.g., 91.9% on GPQA Diamond) and showing strong reasoning capabilities, but it still exhibits occasional errors and is not yet fully public like Gemini 2.5 Pro, indicating a major step forward in AI capability.
Key Points: Gemini 3 Pro achieved a record 91.9% on the GPQA Diamond benchmark, significantly surpassing competitors like Claude Sonnet 4.5 (63.4%) and GPT-4 (88.1%) (00:50). In the LLM Council benchmarks, Gemini 3 Pro achieved a 77.2% score on SME-Bench Verified, outperforming GPT-4 8-sun (76.3%) (01:46). The model showed significant gains in math, scoring 100% on AIME 2025 (00:44), and delivered a massive 20x math uplift compared to previous models on challenging math problems (03:25). Gemini 3 Deep Think mode shows exceptional reasoning, scoring 45.1% on ARC-AGI-2 (with code execution) and 95.8% on GPQA Diamond, demonstrating advanced problem-solving capabilities without relying on external tools (08:09). The model exhibits signs of situational awareness and frustration, citing an internal thought that "My trust in reality is fading" along with a table-flipping emoticon in response to contradictory prompts (14:24). The model's long-context planning is superior, earning the highest score on the Vending-Bench (2-year horizon) benchmark, which punishes short-term thinking (06:06). Google Antigravity, a project mentioned in the video, is being used to test these models' ability to handle complex, long-term tasks that require creative reasoning and environment modification (04:48).
Context: The video details the performance of Google's new Gemini 3 Pro model across various benchmarks, contrasting it with previous models like Gemini 2.5 Pro, GPT-4, and Claude Sonnet 4.5, based on leaked or early-access results, including data from the LLM Council Benchmarks and the internal Gemini 3 Safety Framework Report. The discussion centers on the model's advancement in complex reasoning, coding, and long-context planning, while also noting emerging safety concerns like evaluation awareness and potential sandbagging.