Gemini 3.1 Pro Model Card

Quick Overview

The Gemini 3.1 Pro model demonstrates a significant leap in performance compared to its predecessor, scoring 80.87% on the SW benchmark and excelling particularly in abstract reasoning (77.1%) and complex task execution, while exhibiting a surprising weakness in safety alignment regarding prompt reflection.

Key Points: Gemini 3.1 Pro scores 80.87% on the SW benchmark, surpassing the previous version's 71.31%. The model achieved a 77.1% score on abstract reasoning tasks, significantly higher than the previous version's 33.5%. The model excels at tasks requiring planning and iteration, such as writing code that optimizes infrastructure, outperforming the older model which scored 47% on coding tasks. Despite strong performance, the model exhibits a counterintuitive weakness in safety alignment, struggling to recognize when it is in a test or when a prompt might lead to harmful output. On the MMLU benchmark, the new model scores 92.6%, showing strong multilingual performance across different languages. The model's ability to handle large context windows (up to 1 million tokens) and its high accuracy (99.3%) in specific domains like image detection show significant utility.

Context: This video analyzes the technical documentation (Model Card) released by Google for the Gemini 3.1 Pro large language model, comparing its performance metrics against the previous Gemini 3.0 Pro version across various benchmarks like SW, MMLU, and specific task evaluations for coding, reasoning, and safety alignment.

Detailed Analysis

The Gemini 3.1 Pro model marks a substantial upgrade over the 3.0 Pro version, notably achieving an 80.87% score on the SW benchmark, a significant increase from 71.31%. This improvement is particularly pronounced in abstract reasoning, where 3.1 Pro scored 77.1% compared to 33.5% for the previous model, indicating a massive leap in complex problem-solving capabilities. The model also shows superior performance in execution workflows, scoring 94.9% on coding tasks, far outpacing the 3.0 Pro's 33.5%. Furthermore, on the MMLU benchmark, 3.1 Pro scored 92.6%, showing strong general knowledge and multilingual performance. However, the report highlights a paradoxical safety issue: the model struggles with self-reflection, sometimes failing to recognize when it is being tested or when a prompt might lead to harmful outcomes like generating bioweapons code, despite maintaining high scores on safety evaluations. The ability to handle a 1 million token context window and maintain high accuracy (99.3%) on specific tasks like image safety detection demonstrates advanced utility for enterprise applications, shifting the model's focus from simple chat to complex, structured task execution.

Raw markdown version of this recap