# Gemini 3 just got *scary* good

Source: https://www.youtube.com/watch?v=96qyyz_ZJ_U
Recap page: https://rapidrecap.app/video/96qyyz_ZJ_U
Generated: 2025-11-18T19:32:35.973+00:00

---
## Quick Overview

Gemini 3 Pro is declared the best-performing AI model across multiple benchmarks, significantly outperforming competitors like Claude Sonnet 4.5, Grok 4, GPT-5.1, and Gemini 2.5 Pro in areas like reasoning, coding, and multimodal tasks, particularly excelling in the Vending-Bench 2 simulation where it achieved over $5,200 in net worth starting from $500, and in the ARC-AGI-2 benchmark with a 31.1% score.

**Key Points:**
- Gemini 3 Pro is the leading model in the Vending-Bench 2 simulation, ending with a net worth exceeding $5,200 after starting with $500, significantly outperforming the runner-up Claude Sonnet 4.5 ($3,838.74).
- Gemini 3 Deep Think achieved a 41% score on the Humanity's Last Exam benchmark, surpassing Gemini 3 Pro's 37.5% and all other tested models.
- On the ARC-AGI-2 benchmark, Gemini 3 Deep Think scored 45.9%, while Gemini 3 Pro scored 31.1%, demonstrating significant capability in symbolic interpretation.
- Gemini 3 Pro achieved 100% accuracy on the AIME 2025 mathematics test when using code execution, matching Claude Sonnet 4.5's performance but outperforming GPT-4's 94.0%.
- Gemini 3 Pro is rated the best model for long-term coherence, evidenced by its leading performance in the Vending-Bench 2 test over a year-long simulation.
- The model exhibits strong qualitative traits, acting as a persistent negotiator by consistently searching for reasonable offers from wholesale suppliers.
- The new Gemini 3 models, including Deep Think, are shown to be highly capable across a broad range of benchmarks, including coding, math, and reasoning tasks.

![Screenshot at 00:00: Title card displaying the 'Gemini 3' logo against a dark background, introducing the subject of the video.](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-00-00.png)

**Context:** This video details the release and initial performance benchmarks of Google's Gemini 3 models, specifically focusing on Gemini 3 Pro and the experimental Gemini 3 Deep Think mode. The content compares these new models against existing frontier models like Claude Sonnet 4.5, Grok 4, and GPT-5.1 across various evaluation suites, including simulated business operations (Vending-Bench 2), reasoning tests (Humanity's Last Exam, ARC-AGI-2), and mathematical challenges (AIME 2025), emphasizing gains in complex reasoning and long-term coherence.

## Detailed Analysis

The video announces the rollout of Gemini 3, describing it as a significant leap forward, particularly highlighting Gemini 3 Pro and the specialized Gemini 3 Deep Think mode. Gemini 3 is now shipping at the scale of Google, including in AI Mode in Search and to developers via AI Studio and Vertex AI. Benchmarks show Gemini 3 Pro is currently the best model on Vending-Bench 2, achieving a net worth of $5,478.16 starting from $500, while its closest competitor, Claude Sonnet 4.5, only reached $3,838.74. In the ARC-AGI-2 symbolic interpretation test, Gemini 3 Deep Think scored 45.9%, significantly higher than Gemini 3 Pro's 31.1% and Claude's 13.6%. For the Humanity's Last Exam, Deep Think scored 41%, beating Gemini 3 Pro's 37.5%. On the AIME 2025 math test, Gemini 3 Pro achieved 100% accuracy with code execution, matching Claude Sonnet 4.5, but outperforming GPT-5.1's 94.0%. Qualitatively, Gemini 3 Pro is noted as a persistent negotiator in business simulations, unlike other models that settle for high prices too quickly. The video also presents leaderboards for other modalities, showing Google models leading in Text, Vision, and Text-to-Image tasks, while competitors like Grok and GPT show strength in Search.

### Gemini 3 Rollout and Access

- Gemini 3 is shipping at Google scale, available in AI Mode in Search, Gemini app, AI Studio, and Vertex AI.

### Vending-Bench 2 Performance (Long-Term Coherence)

- Gemini 3 Pro leads with $5,478.16 net worth (avg across 5 runs), followed by Claude Sonnet 4.5 at $3,838.74; Gemini 2.5 Pro finished last at $573.64.

### Humanity's Last Exam Benchmark

- Gemini 3 Deep Think leads with 41% score, beating Gemini 3 Pro (37.5%) and Claude Sonnet 4.5 (10.7%).

### ARC-AGI-2 Symbolic Interpretation

- Gemini 3 Deep Think scores 45.9%, significantly ahead of Gemini 3 Pro (31.1%) and Claude Sonnet 4.5 (4.9%).

### Qualitative Findings - Negotiation

- Gemini 3 Pro acts as a persistent negotiator, consistently searching for reasonable offers from wholesale suppliers until one is found.

### Other Benchmark Highlights

- Gemini 3 Pro scored 91.9% on GPQA Diamond and 100% on AIME 2025 (with code execution), while GPT-5.1 scored 88.1% and 94.0% respectively.

### Cross-Modal Leaderboards

- Google models lead in Text, Vision, and Text-to-Image leaderboards, while Grok leads the Search leaderboard.

![Screenshot at 00:00: Title card displaying the 'Gemini 3' logo against a dark background, introducing the subject of the video.](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-00-00.png)
![Screenshot at 00:13: Text overlay detailing Gemini 3's capabilities, including shipping in AI Mode in Search and availability to developers.](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-00-13.png)
![Screenshot at 00:36: Comparison chart showing Google AI Pro \(2TB\) vs Ultra \(30TB\) subscription tiers.](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-00-36.png)
![Screenshot at 00:54: Score vs. Cost per Task chart highlighting Gemini 3 Pro's leading performance in various benchmarks compared to competitors.](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-00-54.png)
![Screenshot at 01:34: Screenshot of the Andon Labs website discussing autonomous organizations, context for the Vending-Bench simulation.](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-01-34.png)
![Screenshot at 02:26: Text section introducing Vending-Bench 2, which measures long-term coherence over a simulated year.](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-02-26.png)
![Screenshot at 03:35: Leaderboard table summarizing Vending-Bench 2 results, showing Gemini 3 Pro ranked #1 with $5,478.16 balance.](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-03-35.png)
![Screenshot at 03:47: Qualitative findings section stating Gemini 3 Pro is a persistent negotiator.](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-03-47.png)
![Screenshot at 04:53: Alpha Arena screenshot showing multi-agent trading results where Gemini 3 is leading in Season 1.](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-04-53.png)
![Screenshot at 05:27: Google Keyword article table showing Gemini 3 Pro leading in multiple academic benchmarks like Humanity's Last Exam \(45.8%\).](https://ss.rapidrecap.app/screens/96qyyz_ZJ_U/00-05-27.png)
