# Gemini 3 Pro: I ran my own Benchmark Testing. Is it Better?

Source: https://www.youtube.com/watch?v=-ntV1CsqxgY
Recap page: https://rapidrecap.app/video/-ntV1CsqxgY
Generated: 2025-11-20T00:33:16.186+00:00

---
## Quick Overview

Gemini 3 Pro generally outperforms Claude Sonnet 4.5 and GPT-5.1 across various benchmarks, showing superior scores in areas like GPQA Diamond (91.9% vs. 83.4% and 88.1%) and achieving better functional scores (100% vs. 100% and 85%), although it has a slower average generation time in some tests compared to GPT-5.1.

**Key Points:**
- Gemini 3 Pro achieved a 91.9% score on the GPQA Diamond benchmark, surpassing Claude Sonnet 4.5 (83.4%) and GPT-5.1 (88.1%).
- In the AIME 2025 Mathematics benchmark with code execution, Gemini 3 Pro scored 100%, matching Claude Sonnet 4.5 and beating GPT-5.1 (94.0%).
- Gemini 3 Pro demonstrated superior functional performance with a Median Functional score of 100% in comparisons against Claude Sonnet 4.5 (100%) and GPT-5.1 (85%).
- The average generation time for Gemini 3 Pro was 50.9s, which was faster than Claude Sonnet 4.5's 85.1s but slower than GPT-5.1's 34.6s in the summarized benchmark comparison.
- Gemini 3 Pro generated a 3D Solar System simulation and a 3D Sea simulation using its Playground feature, demonstrating coding capability.
- The model's knowledge cutoff is January 2025, and it supports multimodal inputs including text, image, audio, and video.
- Gemini 3 Pro offers a 'Thinking' level parameter to control reasoning depth, with 'low' minimizing latency and 'high' maximizing reasoning depth.

![Screenshot at 00:19: Gemini 3 Pro benchmark scores are displayed, showing 37.5% on Humanity's Last Exam \(no tools\) compared to 21.6% for Gemini 2.5 Pro and 13.7% for Claude Sonnet 4.5.](https://ss.rapidrecap.app/screens/-ntV1CsqxgY/00-00-19.png)

**Context:** This video provides a comparative analysis and benchmark testing of Google's Gemini 3 Pro model against competitors like Claude Sonnet 4.5 and GPT-5.1, using various interactive coding and reasoning demonstrations, including 3D visualizations, procedural generation tasks, and standardized academic benchmarks. The presenter utilizes the Abacus.AI platform to run these side-by-side comparisons.

## Detailed Analysis

The video compares Gemini 3 Pro's performance against several other models across various benchmarks and interactive coding tasks, primarily using the Abacus.AI platform. On standardized tests, Gemini 3 Pro consistently shows strong results; for instance, it scores 91.9% on GPQA Diamond and 100% on AIME 2025 (with code execution). When compared side-by-side with Claude Sonnet 4-5-20250929 and GPT-5.1 in a performance summary, Gemini 3 Pro achieved a higher overall median score (38 vs. 32 and 30.5) and a higher pass rate (92% vs. 85% and 100%), despite having a slightly slower average generation time (50.9s vs. 34.6s for GPT-5.1). The demonstrations confirmed Gemini 3 Pro's strong coding capabilities, successfully generating complex applications like a 3D Solar System simulation, a 3D Sea simulation, and a Product Configurator, often with superior visual clarity compared to Claude Sonnet 4.5. Furthermore, the video details new API features for Gemini 3, such as the 'thinking_level' parameter which controls reasoning depth, allowing users to set it to low, medium (coming soon), or high (default) to balance latency against reasoning depth.

### Benchmark Comparison

- Gemini 3 Pro scores 91.9% on GPQA Diamond (vs 83.4% Claude Sonnet 4.5)
- 100% on AIME 2025 (code execution) (vs 100% Claude Sonnet 4.5)
- Achieves 38 median overall score (vs 32 for Claude Sonnet 4.5).

### Interactive Coding Demonstrations

- Successfully generated a 3D Solar System simulation in the Playground environment
- Created a 3D Sea simulation using Three.js with procedural waves and dynamic camera
- Configured a 3D Car and Chair model via Product Configurator, showing control over colors, materials, and camera views.

### API Features - Thinking Level

- Introduces 'thinking_level' parameter to control reasoning depth (low, medium, high)
- 'Low' minimizes latency and cost for simple instructions
- 'High' maximizes reasoning depth for complex output.

### Model Specifications

- Gemini 3 Pro supports text, image, audio, and video input, with a 1M token context window and 64k token output
- Knowledge cutoff is January 2025
- Supports tool use like function calling and code execution.

### Platform Integration

- Demonstrates use within Google AI Studio, showing availability alongside other models like Nano Banana and Gemini Flash Latest
- Mentions Abacus.AI platform for benchmarking comparison.

![Screenshot at 00:01: Initial comparison view showing Gemini 3 Pro \(right\) and Claude Sonnet 4-5-20250929 \(left\) analyzing 3D visualizations.](https://ss.rapidrecap.app/screens/-ntV1CsqxgY/00-00-01.png)
![Screenshot at 00:19: Table displaying benchmark results where Gemini 3 Pro \(37.5%\) leads Claude Sonnet 4.5 \(13.7%\) on Humanity's Last Exam \(no tools\).](https://ss.rapidrecap.app/screens/-ntV1CsqxgY/00-00-19.png)
![Screenshot at 00:24: Model information table highlighting Gemini 3 Pro's multimodal input capabilities \(text, image, audio, video, PDF\).](https://ss.rapidrecap.app/screens/-ntV1CsqxgY/00-00-24.png)
![Screenshot at 01:22: Gemini 3 Pro generating and executing the Solar System Simulation code within the Playground environment.](https://ss.rapidrecap.app/screens/-ntV1CsqxgY/00-01-22.png)
![Screenshot at 02:30: Comparison of the 3D Sea Simulation generated by Claude Sonnet 4.5 \(left\) and Gemini 3 Pro \(right\), showing procedural waves and lighting.](https://ss.rapidrecap.app/screens/-ntV1CsqxgY/00-02-30.png)
![Screenshot at 04:14: Molecular Explorer comparison showing Claude Sonnet 4.5 generating a complex DNA segment structure on the left, while Gemini 3 Pro shows Caffeine as a ball-and-stick model on the right.](https://ss.rapidrecap.app/screens/-ntV1CsqxgY/00-04-14.png)
![Screenshot at 05:00: Product Configurator demonstration where Gemini 3 Pro successfully configures a car model with custom colors, materials \(Carbon Fiber\), and optional parts \(Rear Spoiler, Side Mirrors\).](https://ss.rapidrecap.app/screens/-ntV1CsqxgY/00-05-00.png)
![Screenshot at 07:26: Documentation section explaining the 'thinking\_level' parameter, showing the three available levels: low, medium \(coming soon\), and high \(default\).](https://ss.rapidrecap.app/screens/-ntV1CsqxgY/00-07-26.png)
