Gemini 3 Pro: I ran my own Benchmark Testing. Is it Better?

Quick Overview

Gemini 3 Pro generally outperforms Claude Sonnet 4.5 and GPT-5.1 across various benchmarks, showing superior scores in areas like GPQA Diamond (91.9% vs. 83.4% and 88.1%) and achieving better functional scores (100% vs. 100% and 85%), although it has a slower average generation time in some tests compared to GPT-5.1.

Key Points: Gemini 3 Pro achieved a 91.9% score on the GPQA Diamond benchmark, surpassing Claude Sonnet 4.5 (83.4%) and GPT-5.1 (88.1%). In the AIME 2025 Mathematics benchmark with code execution, Gemini 3 Pro scored 100%, matching Claude Sonnet 4.5 and beating GPT-5.1 (94.0%). Gemini 3 Pro demonstrated superior functional performance with a Median Functional score of 100% in comparisons against Claude Sonnet 4.5 (100%) and GPT-5.1 (85%). The average generation time for Gemini 3 Pro was 50.9s, which was faster than Claude Sonnet 4.5's 85.1s but slower than GPT-5.1's 34.6s in the summarized benchmark comparison. Gemini 3 Pro generated a 3D Solar System simulation and a 3D Sea simulation using its Playground feature, demonstrating coding capability. The model's knowledge cutoff is January 2025, and it supports multimodal inputs including text, image, audio, and video. Gemini 3 Pro offers a 'Thinking' level parameter to control reasoning depth, with 'low' minimizing latency and 'high' maximizing reasoning depth.

Context: This video provides a comparative analysis and benchmark testing of Google's Gemini 3 Pro model against competitors like Claude Sonnet 4.5 and GPT-5.1, using various interactive coding and reasoning demonstrations, including 3D visualizations, procedural generation tasks, and standardized academic benchmarks. The presenter utilizes the Abacus.AI platform to run these side-by-side comparisons.

Detailed Analysis

The video compares Gemini 3 Pro's performance against several other models across various benchmarks and interactive coding tasks, primarily using the Abacus.AI platform. On standardized tests, Gemini 3 Pro consistently shows strong results; for instance, it scores 91.9% on GPQA Diamond and 100% on AIME 2025 (with code execution). When compared side-by-side with Claude Sonnet 4-5-20250929 and GPT-5.1 in a performance summary, Gemini 3 Pro achieved a higher overall median score (38 vs. 32 and 30.5) and a higher pass rate (92% vs. 85% and 100%), despite having a slightly slower average generation time (50.9s vs. 34.6s for GPT-5.1). The demonstrations confirmed Gemini 3 Pro's strong coding capabilities, successfully generating complex applications like a 3D Solar System simulation, a 3D Sea simulation, and a Product Configurator, often with superior visual clarity compared to Claude Sonnet 4.5. Furthermore, the video details new API features for Gemini 3, such as the 'thinkinglevel' parameter which controls reasoning depth, allowing users to set it to low, medium (coming soon), or high (default) to balance latency against reasoning depth.

Raw markdown version of this recap