# Here's What They Didn't Tell You About Gemini 3

Source: https://www.youtube.com/watch?v=kyflIo3EKLw
Recap page: https://rapidrecap.app/video/kyflIo3EKLw
Generated: 2025-11-19T14:36:18.974+00:00

---
## Quick Overview

Gemini 3 outperforms Claude Sonnet 4.5 and GPT-5.1 on coding benchmarks like LiveCodeBench Pro and SWE-Bench, demonstrating superior reasoning and agentic capabilities, although the CLI implementation for Gemini 3 caused initial frustration due to missing features like sound effects and required multiple reprompts to fix.

**Key Points:**
- Gemini 3 significantly beats Claude Sonnet 4.5 by nearly 1,000 Elo points and GPT-5.1 by 200 points on the LiveCodeBench Pro competitive programming benchmark (Elo 2,439).
- On the SWE-Bench Verified benchmark (real-world GitHub issues), Gemini 3 Pro, Claude 4.5, and GPT-5.1 are in a close performance tie, all scoring between 76-77% (a massive jump from the ~60% tier).
- Gemini 3 showed a large gap on Terminal-Bench 2.0 (Agentic Terminal Coding), scoring 54.2% against Claude's 42.8% and GPT-5.1's 47.6%.
- The initial implementation of Gemini 3 via the CLI was frustrating, requiring about 20 reprompts to fix errors that Claude only needed about 4 to resolve.
- The new Agentic development platform, Google Antigravity, was released alongside Gemini 3, enabling agents to autonomously plan and execute complex, end-to-end software tasks.
- The MonkeyType clone app, built using the Gemini 3 agent, featured a highly polished UI with smooth scrolling and custom themes, though the initial sound effects were not implemented correctly.

![Screenshot at 0:05: The pinned tweet from the official GeminiApp account announcing Gemini 3, highlighting its three core components: the intelligent model, generative interfaces, and the Gemini Agent.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-00-05.png)

**Context:** This video reviews and benchmarks the newly announced Google Gemini 3 model against competitors like Anthropic's Claude Sonnet 4.5 and OpenAI's GPT-5.1, focusing heavily on coding and agentic performance metrics. The creator also tests the new Google Antigravity agentic development platform by having it build a complex web-based macOS clone application, comparing the results and implementation experience between Gemini 3 and Claude.

## Detailed Analysis

The video analyzes the official announcement of Google's Gemini 3 model, comparing its performance across several coding and agentic benchmarks against Claude Sonnet 4.5 and GPT-5.1. In competitive programming (LiveCodeBench Pro), Gemini 3 achieved a massive 2,439 Elo score, beating Claude by nearly 1,000 points and GPT-5.1 by 200 points. On SWE-Bench (real-world GitHub issues), all three leading models are neck-and-neck at 76-77%. Gemini 3 demonstrated a significant lead in agentic terminal coding (Terminal-Bench 2.0) with 54.2%. The creator tested Gemini 3's agent capabilities by having it build a complex 'WebOS-macOS Clone' application using the new Antigravity platform. While the resulting UI was deemed excellent and smooth, the initial experience using the Gemini CLI was frustrating, requiring about 20 reprompts to fix issues, compared to Claude which fixed similar issues in about 4 prompts. The model card confirms Gemini 3 features a massive 1M token context window and native multimodality, but the smaller context window variant was not referenced in the documentation.

### Gemini 3 Official Announcement

- Gemini 3 is announced as Google's most intelligent model, featuring generative interfaces and the Gemini Agent; it is shipping at the scale of Google, including in AI Mode in Search and to developers via AI Studio and Vertex AI.

### Coding Benchmark Performance

- Gemini 3 beats Claude Sonnet 4.5 by nearly 1,000 Elo points on LiveCodeBench Pro; all three top models tie at 76-77% on SWE-Bench; Gemini 3 leads GPT-5.1 and Claude on Terminal-Bench 2.0 (54.2% vs 47.6% and 42.8%).

### Agentic Development with Antigravity

- The new agent-first platform, Google Antigravity, allows agents to autonomously plan and execute complex tasks; the creator used it to build a functional macOS clone app using Gemini 3 Pro.

### Agent Performance Comparison

- Claude required about 4 reprompts to fix issues during the macOS clone build, while Gemini 3 required about 20 reprompts, although Gemini 3 finished the entire task 20 minutes faster due to its speed.

### MonkeyType Clone App Testing

- The AI-generated typing test app featured highly polished UI elements, including smooth scrolling and custom themes, but initial testing revealed missing sound effects and UI rendering issues that required prompting to fix.

![Screenshot at 0:05: The pinned tweet from the official GeminiApp account announcing Gemini 3, highlighting its three core components: the intelligent model, generative interfaces, and the Gemini Agent.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-00-05.png)
![Screenshot at 0:24: A comparison chart showing Claude Sonnet 4 vs. Kimi K2 0905 pricing structure for input and output tokens.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-00-24.png)
![Screenshot at 0:41: Highlighting the distribution of Gemini 3 across Google products, including AI Mode in Search, the Gemini app, AI Studio, Vertex AI, and the Google Antigravity platform.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-00-41.png)
![Screenshot at 1:03: The agent's breakdown of coding benchmarks, showing Gemini 3 Pro beats Claude Sonnet 4.5 by nearly 1,000 Elo points on LiveCodeBench Pro.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-01-03.png)
![Screenshot at 2:24: The Antigravity agent interface executing the first step of scaffolding the 'Web-based macOS Clone' project.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-02-24.png)
![Screenshot at 3:12: The Gemini CLI model selection menu, showing Gemini 3 Pro \(preview\) is the default choice for complex tasks.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-03-12.png)
![Screenshot at 4:25: The Settings application within the WebOS clone, showing the prompt to 'Generate with Nano Banana \(AI\)', which is the codename for Gemini 2.5 Flash Image model.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-04-25.png)
![Screenshot at 6:28: A summary slide comparing Claude 4.5 and Gemini 3 testing results on Story 7, noting Gemini 3 used only 4% context and finished 20 minutes faster.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-06-28.png)
![Screenshot at 7:54: The user story/requirements document for Phase 1 of the MonkeyType clone, detailing required features like smooth scrolling, keyboard shortcuts, and sound effects.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-07-54.png)
![Screenshot at 8:11: The final MonkeyType clone application running, showcasing the polished UI implementation achieved by the AI agent.](https://ss.rapidrecap.app/screens/kyflIo3EKLw/00-08-11.png)
