# 🎅 ThursdAI - December 18 - Gemini 3 Flash, Nemotron 3 Nano, ChatGPT Image 1.5 & More AI

Source: https://www.youtube.com/watch?v=eQeVAYQYBGs
Recap page: https://rapidrecap.app/video/eQeVAYQYBGs
Generated: 2025-12-19T00:05:41.56+00:00

---
## Quick Overview

The ThursdAI episode for December 18, 2025, covered major AI releases including Google's Gemini 3 Flash, which demonstrated superior price-to-intelligence ratio and multimodal capabilities, NVIDIA's open-source Nemotron 3 Nano model showing significant efficiency gains, OpenAI's GPT Image 1.5 achieving massive speed/cost improvements and leading on LMSYS Arena, xAI's Grok Voice Agent API ranking #1 in Big Bench Audio, and Meta's release of SAM Audio for zero-shot audio source separation.

**Key Points:**
- Google launched Gemini 3 Flash, a frontier-class model offering Gemini 3 Pro-level reasoning at significantly lower cost ($0.50/1M input tokens) and high speed, impacting Google Search and Assistant.
- NVIDIA released Nemotron 3 Nano, an open-source 30B parameter model using a hybrid Mamba-MoE architecture, achieving 1.5-3.3x faster inference than comparable models while maintaining high accuracy.
- OpenAI released GPT Image 1.5, which is 4x faster and 20% cheaper than previous versions, topping the LMSYS Image Arena leaderboards and showing strong performance in preserving subject details.
- xAI released its Grok Voice Agent API, claiming the #1 spot on the Big Bench Audio benchmark with 92.3% accuracy and offering 5 cents per minute flat rate, powering Tesla vehicles.
- Meta released SAM Audio, a unified multimodal model for zero-shot audio source separation using text, visual, or temporal prompts, capable of isolating specific sounds like music or a train.
- The episode highlighted the rapid acceleration in AI, particularly noting that Google's releases this week were not reliant on the slow, iterative approach seen in previous years.
- The hosts also covered the release of FunctionGemma, a specialized 270M parameter model for function calling designed to run on edge devices.

![Screenshot at 00:03: 03:Wolfram Ravenwolf highlights the announcement of NVIDIA's Nemotron 3 Nano, a 30B parameter, 3.5B active parameter hybrid Mamba-MoE model, emphasizing its open-source nature and 1.5-3.3x faster inference.](https://ss.rapidrecap.app/screens/eQeVAYQYBGs/00-00-03.png)

**Context:** This episode of ThursdAI, hosted by Alex Volkov and featuring guests like Wolfram Ravenwolf, Yam Peleg, Ryan Carson, Kwindla Hultman Kramer, nisten, and LDJ, provided a year-end recap of significant AI releases around December 18, 2025. The discussion centered on major announcements from Google, NVIDIA, OpenAI, xAI, and Meta, focusing heavily on performance metrics, cost-efficiency, and the increasing trend toward open-source contributions and multimodal capabilities.

## Detailed Analysis

The ThursdAI episode covered several major AI announcements from the week ending December 18, 2025. Alex Volkov kicked off by noting the fast pace of development, especially from Google. Google released Gemini 3 Flash, a frontier-class model focused on speed and cost, priced at $0.50/1M input tokens and $3.00/1M output tokens. It performs comparably to Gemini 3 Pro on many benchmarks while being significantly faster and cheaper, powering services like Google Search and the Gemini Assistant. The speaker noted that this speed/intelligence ratio is 'absolutely ridiculous.' Wolfram Ravenwolf highlighted NVIDIA's release of Nemotron 3 Nano, an open-source hybrid Mamba-MoE model with 30B total parameters (only 3.5B active), emphasizing NVIDIA's commitment to open weights/data sets and its superior performance on coding benchmarks like SWE-Bench Pro (78%) compared to other models. The discussion moved to multimodal advancements, noting xAI's Grok Voice Agent API ranked #1 on Big Bench Audio (92.3% accuracy in speech-to-speech reasoning) at a low cost ($0.05/min flat rate) and is powering Tesla vehicles. Meta released SAM Audio, a unified multimodal model for zero-shot audio source separation using text, visual, or temporal prompts, which can isolate specific sounds like drums from music. OpenAI also released GPT Image 1.5, which is 4x faster and 20% cheaper than its predecessor, showing high performance in image editing benchmarks, though some noted it takes more creative liberties than Nano Banana. Finally, Google released FunctionGemma, a small, specialized 270M parameter model for function calling that runs on edge devices like phones and browsers, which is seen as a major step for agentic AI.

### Major LLM & Model Releases

- Google launched Gemini 3 Flash (frontier-class, fast, cheap, multimodal)
- NVIDIA released Nemotron 3 Nano (open-source, hybrid Mamba-MoE, 1.5-3.3x faster inference)
- OpenAI launched GPT Image 1.5 (4x faster, 20% cheaper, leading LMSYS Arena)
- Google announced FunctionGemma (270M parameter, edge-optimized function calling model)

### Voice & Audio Updates

- xAI launched Grok Voice Agent API (#1 on Big Bench Audio with 92.3% accuracy)
- Resemble AI released Chatterbox Turbo (open-source TTS beating ElevenLabs/Cartesia)
- Meta released SAM Audio (unified zero-shot audio source separation)

### Key Performance Metrics (Gemini 3 Flash)

- $0.50/1M input tokens, $3.00/1M output tokens
- 100 simultaneous function calls
- 100% accuracy on Mathematics (no tools)

### Key Performance Metrics (Nemotron 3 Nano)

- 1.5-3.3x faster inference than Qwen-3-30B-A3B
- 99.7% on AIME 2025 with code execution

### Key Performance Metrics (Grok Voice Agent API)

- 92.3% accuracy on speech-to-speech reasoning
- $0.05/min flat rate, nearly 5x faster than OpenAI ($0.10/min)

### Key Performance Metrics (GPT Image 1.5)

- 4x faster generation, 20% cheaper, topped LMSYS Image Arena leaderboards (1264 ELO for text-to-image)

### Open Source Focus

- NVIDIA's Nemotron 3 Nano release emphasized open weights/data sets, contrasting with the more closed approach seen in some other major releases.

![Screenshot at 00:00: Alex Volkov opens the ThursdAI episode wearing a Santa hat, welcoming viewers to the December 18, 2025 broadcast.](https://ss.rapidrecap.app/screens/eQeVAYQYBGs/00-00-00.png)
![Screenshot at 01:03: The TL;DR slide for the episode lists topics including Gemini 3 Flash, Nemotron 3 Nano, Grok Voice, SAM Audio, and Chatterbox.](https://ss.rapidrecap.app/screens/eQeVAYQYBGs/00-01-03.png)
![Screenshot at 03:40: An infographic detailing NVIDIA's Nemotron 3 Nano, highlighting its 30B parameters, 3.5B active, hybrid Mamba-MoE architecture, and 1.5-3.3x faster inference.](https://ss.rapidrecap.app/screens/eQeVAYQYBGs/00-03-40.png)
![Screenshot at 05:54: A multi-panel video conference view of all the hosts and guests participating in the weekly AI discussion.](https://ss.rapidrecap.app/screens/eQeVAYQYBGs/00-05-54.png)
![Screenshot at 07:04: A slide showing Bolmo 7B vs. byte-level and subword peers, illustrating the performance gains of byte-level models.](https://ss.rapidrecap.app/screens/eQeVAYQYBGs/00-07-04.png)
