# 1,000 tok/s?! The Age of Diffusion Based LLMs Is Upon Us

Source: https://www.youtube.com/watch?v=TVlFf_Po1bs
Recap page: https://rapidrecap.app/video/TVlFf_Po1bs
Generated: 2025-07-28T15:01:59.829+00:00

---
## Quick Overview

Diffusion-based Large Language Models (LLMs) are emerging as a competitive alternative to traditional autoregressive LLMs, particularly in coding and multimodal tasks, demonstrating faster generation speeds and comparable or superior performance on various benchmarks.

**Key Points:**
- Diffusion-based LLMs, like Inception AI's Mercury and Google DeepMind's Gemini Diffusion, are becoming competitive alternatives to autoregressive LLMs.
- Diffusion LLMs generate text by iteratively refining noisy outputs, allowing for more flexible and potentially better performance on tasks with non-sequential dependencies.
- Mercury Coder Mini and Small models offer significant speed advantages (over 5x faster) and strong coding performance compared to other models.
- Gemini Diffusion demonstrates state-of-the-art text diffusion capabilities, excelling in coding and math tasks due to its iterative refinement process.
- Benchmarks show diffusion LLMs generally outperform autoregressive models on coding tasks and are competitive in multimodal understanding.
- Research papers like LaViDA and MMADa showcase the potential of diffusion models for multimodal tasks, integrating vision and language understanding.
- The core advantage of diffusion LLMs lies in their ability to handle tasks requiring non-sequential or long-range dependencies, offering a new paradigm in AI development.

![Screenshot at 00:00: Title card and introduction to diffusion LLMs, setting the stage for the comparison between diffusion and autoregressive models.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-00-00.png)

**Context:** The video delves into the rapidly evolving field of Large Language Models (LLMs), specifically focusing on the emergence and capabilities of diffusion-based LLMs. It contrasts these new models with the established autoregressive LLMs, highlighting the potential advantages of diffusion models in terms of speed, flexibility, and performance on complex tasks, particularly in coding and multimodal applications.

## Detailed Analysis

This video explores the rise of diffusion-based Large Language Models (LLMs) as a new paradigm in AI, contrasting them with traditional autoregressive LLMs. It highlights how diffusion LLMs, like Inception AI's Mercury models and Google DeepMind's Gemini Diffusion, operate by progressively refining noisy outputs over multiple steps, which allows for more flexible generation and potentially better performance on complex tasks. The video showcases Mercury Coder Mini and Small as coding-focused models that are significantly faster than existing models like GPT-4o and Claude 3.5 Haiku, while also achieving competitive results on coding benchmarks. Gemini Diffusion is presented as a state-of-the-art text diffusion model that excels at coding and math due to its iterative refinement process. Benchmarks are presented comparing diffusion LLMs against autoregressive models, showing that diffusion models generally perform better on coding tasks and are competitive in multimodal understanding tasks, though they lag slightly in reasoning and multilingual benchmarks. The video also touches upon the potential of diffusion LLMs for tasks like in-filling and early coordination, and mentions research papers like LaViDA and MMADa as examples of diffusion-based multimodal models. The core advantage of diffusion LLMs lies in their ability to handle tasks that require non-sequential or long-range dependencies, which is difficult for autoregressive models.

### Introduction to Diffusion LLMs

- Overview of diffusion-based LLMs, contrasting them with autoregressive LLMs
- Explanation of the iterative refinement process
- Mention of Inception AI's Mercury models and Google DeepMind's Gemini Diffusion

### Performance Comparison

- Benchmarking diffusion LLMs against autoregressive models
- Mercury Coder Mini and Small speed and performance metrics
- Gemini Diffusion speed and benchmark results
- Discussion of coding vs. non-coding benchmarks

### Key Advantages of Diffusion LLMs

- Ability to handle tasks with non-sequential or long-range dependencies
- Potential for in-filling and early coordination
- Multimodal capabilities with image understanding and generation

### Research Examples

- LaViDA (Large Vision-Language Diffusion Model with Masking)
- MMADa (Multimodal Diffusion Language Models)
- Discussion of their architectures and performance

![Screenshot at 00:00: Title card and introduction to diffusion LLMs.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-00-00.png)
![Screenshot at 00:44: Bar chart comparing output speed of various LLMs, highlighting Mercury Coder Mini's performance.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-00-44.png)
![Screenshot at 01:03: Tweet announcing Gemini Diffusion, Google DeepMind's text diffusion model.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-01-03.png)
![Screenshot at 01:37: Tweet from Inception AI announcing Mercury, the first commercial-grade diffusion LLM.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-01-37.png)
![Screenshot at 04:43: Diagram illustrating the iterative denoising process in diffusion models.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-04-43.png)
![Screenshot at 07:08: LM Arena leaderboard showing comparative performance of different LLM assistants.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-07-08.png)
![Screenshot at 07:22: Copilot Arena leaderboard showing Mercury Coder Mini's ranking and performance.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-07-22.png)
![Screenshot at 08:03: Benchmark results comparing Gemini Diffusion and Gemini 2.0 Flash-Lite on various tasks.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-08-03.png)
![Screenshot at 10:05: LaViDA: A Large Diffusion Language Model for Multimodal Understanding paper cover.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-10-05.png)
![Screenshot at 11:03: MMADa pipeline overview diagram, showcasing its multimodal capabilities.](https://ss.rapidrecap.app/screens/TVlFf_Po1bs/00-11-03.png)
