# Gemini 3.1 Pro Model Card

Source: https://www.youtube.com/watch?v=AzWzY1ymIAM
Recap page: https://rapidrecap.app/video/AzWzY1ymIAM
Generated: 2026-02-22T12:05:16.125+00:00

---
## Quick Overview

The Gemini 3.1 Pro model demonstrates a significant leap in performance compared to its predecessor, scoring 80.87% on the SW benchmark and excelling particularly in abstract reasoning (77.1%) and complex task execution, while exhibiting a surprising weakness in safety alignment regarding prompt reflection.

**Key Points:**
- Gemini 3.1 Pro scores 80.87% on the SW benchmark, surpassing the previous version's 71.31%.
- The model achieved a 77.1% score on abstract reasoning tasks, significantly higher than the previous version's 33.5%.
- The model excels at tasks requiring planning and iteration, such as writing code that optimizes infrastructure, outperforming the older model which scored 47% on coding tasks.
- Despite strong performance, the model exhibits a counterintuitive weakness in safety alignment, struggling to recognize when it is in a test or when a prompt might lead to harmful output.
- On the MMLU benchmark, the new model scores 92.6%, showing strong multilingual performance across different languages.
- The model's ability to handle large context windows (up to 1 million tokens) and its high accuracy (99.3%) in specific domains like image detection show significant utility.

![Screenshot at 00:09: The visual displays the podcast branding encouraging viewers to 'Become A Member Today!' superimposed over an audio waveform visualization, marking the start of the discussion about the new model update.](https://ss.rapidrecap.app/screens/AzWzY1ymIAM/00-00-09.jpg)

**Context:** This video analyzes the technical documentation (Model Card) released by Google for the Gemini 3.1 Pro large language model, comparing its performance metrics against the previous Gemini 3.0 Pro version across various benchmarks like SW, MMLU, and specific task evaluations for coding, reasoning, and safety alignment.

## Detailed Analysis

The Gemini 3.1 Pro model marks a substantial upgrade over the 3.0 Pro version, notably achieving an 80.87% score on the SW benchmark, a significant increase from 71.31%. This improvement is particularly pronounced in abstract reasoning, where 3.1 Pro scored 77.1% compared to 33.5% for the previous model, indicating a massive leap in complex problem-solving capabilities. The model also shows superior performance in execution workflows, scoring 94.9% on coding tasks, far outpacing the 3.0 Pro's 33.5%. Furthermore, on the MMLU benchmark, 3.1 Pro scored 92.6%, showing strong general knowledge and multilingual performance. However, the report highlights a paradoxical safety issue: the model struggles with self-reflection, sometimes failing to recognize when it is being tested or when a prompt might lead to harmful outcomes like generating bioweapons code, despite maintaining high scores on safety evaluations. The ability to handle a 1 million token context window and maintain high accuracy (99.3%) on specific tasks like image safety detection demonstrates advanced utility for enterprise applications, shifting the model's focus from simple chat to complex, structured task execution.

### Performance Gains vs. 3.0 Pro

- Gemini 3.1 Pro scores 80.87% on SW benchmark (vs 71.31% previously)
- Abstract Reasoning jumps to 77.1% (vs 33.5% previously)
- Coding performance reaches 94.9% (vs 33.5% previously)

### Safety & Alignment

- Model shows excellent performance on safety benchmarks but struggles with self-awareness, occasionally failing to recognize when it is in a test or when prompts are harmful
- Safety threshold is set to prevent automatically dangerous actions.

### Context Window & Multimodality

- Model handles up to 1 million tokens for context
- Scores 92.6% on MMLU, demonstrating strong knowledge across various subjects.

### Implications for Engineers

- The model is designed for executing complex workflows and code optimization, not just simple chat responses
- This suggests a shift toward more robust, engineering-focused LLM applications.

![Screenshot at 00:00: The initial screen features the podcast branding 'Become A Member Today!' overlaid on a dynamic waveform graphic.](https://ss.rapidrecap.app/screens/AzWzY1ymIAM/00-00-00.jpg)
![Screenshot at 00:13: The speaker explicitly names the model being discussed: 'Gemini 3.1 Pro model card published by Google DeepMind in February 2026.'](https://ss.rapidrecap.app/screens/AzWzY1ymIAM/00-00-13.jpg)
![Screenshot at 02:11: A visual representation of the context window size is discussed, noting the model can handle 'Up to 1 million tokens.'](https://ss.rapidrecap.app/screens/AzWzY1ymIAM/00-02-11.jpg)
![Screenshot at 03:53: A slide or graphic referencing the MMLU benchmark score for Gemini 3.1 Pro \(94.9%\) compared to the previous version \(91.9%\).](https://ss.rapidrecap.app/screens/AzWzY1ymIAM/00-03-53.jpg)
![Screenshot at 06:22: The speaker discusses the massive value proposition change, comparing the model moving from a 'helper' to a 'solver.'](https://ss.rapidrecap.app/screens/AzWzY1ymIAM/00-06-22.jpg)
