Gemini 3 Flash Model Card

Quick Overview

The Gemini 3 Flash model significantly outperforms its predecessor, Gemini 2.5 Pro, by achieving a 10-point higher score on the MMLU benchmark and demonstrating a 75% reduction in input cost while maintaining a 15% improvement in accuracy on hard data extraction tasks compared to the previous model.

Key Points: Gemini 3 Flash achieved a 10-point higher score on the MMLU benchmark compared to Gemini 2.5 Pro. The new model shows a 75% reduction in input cost for complex reasoning tasks, dropping from $2.00 to $0.50 per million tokens. Accuracy on hard data extraction, such as reading complex spreadsheets and handwritten notes, improved by 15% over the prior model. Gemini 3 Flash is natively multimodal, handling text, audio, images, and video simultaneously, unlike previous models that often stitched these modalities together. The model's ability to orchestrate complex, multi-step reasoning across vast datasets is highlighted as a major advancement, enabling real-time strategic advice. The team is confident that the speed and cost-effectiveness of Gemini 3 Flash make advanced AI accessible to smaller businesses and developers. The model's superior performance on complex reasoning and multi-step tasks is evidenced by its ability to score 81% on the big multimodal benchmark, compared to 33% for the previous model.

Context: The video discusses the release and capabilities of Google's new large language model, Gemini 3 Flash, contrasting it with its predecessor, Gemini 2.5 Pro. The hosts, speaking in a podcast format, focus on quantifiable improvements in performance metrics, cost efficiency, and advanced reasoning capabilities, particularly how these advancements translate to practical enterprise applications like data engineering and business strategy.

Detailed Analysis

The discussion centers on the new Gemini 3 Flash model, positioning it as a significant leap over Gemini 2.5 Pro. The primary outcome is that Gemini 3 Flash is substantially faster, cheaper, and more capable, especially in complex, multimodal reasoning. Specifically, it scores 10 points higher on the MMLU benchmark. Cost savings are dramatic: input costs dropped from $2.00 to $0.50 per million tokens, a 75% reduction. Furthermore, accuracy on difficult data extraction (like handwriting and spreadsheets) improved by 15%. The model's high performance is quantified by an 81% score on a large multimodal benchmark, crushing the previous model's 33%. A key differentiator is its native multimodality and ability to orchestrate complex, multi-step reasoning across text, audio, images, and video simultaneously, effectively acting as an autonomous agent that can manage entire workflows from data analysis to marketing, all while staying within safety guardrails.

Raw markdown version of this recap