# High-Efficiency Diffusion Models for On-Device Image Generation and Editing [Hung Bui] - 753

Source: https://www.youtube.com/watch?v=klk9oher8Bo
Recap page: https://rapidrecap.app/video/klk9oher8Bo
Generated: 2025-11-11T19:38:58.909+00:00

---
## Quick Overview

Hung Bui, Qualcomm's VP of Technology, explains that their strategy for efficient on-device AI, particularly for image generation and editing, involves distilling knowledge from large models into smaller, highly efficient models (like their 7-billion parameter model) through distillation techniques, allowing them to perform nearly as well as larger models while running locally on mobile devices, thus addressing privacy and latency concerns.

**Key Points:**
- Qualcomm recently acquired VinAI Research, which ranked in the top 25 industrial AI labs based on research output.
- The team successfully developed a 7-billion parameter model that performs even better than the 70-billion parameter model it was distilled from.
- A key focus is achieving high-quality image generation and editing using these much smaller models, aiming for performance comparable to larger models.
- The research focuses on efficiency, specifically enabling models to run directly on mobile devices to address privacy and latency issues.
- The process involves distilling knowledge from large teacher models (like the 70-billion parameter model) into smaller student models (like the 7-billion parameter model).
- The approach involves running the inference process on the device itself rather than relying on cloud computation.
- The research group in Vietnam, formerly Stanford Research Institute (SRI) International's Vietnam AI Center, has a history dating back to 2019.

![Screenshot at 00:51: Introduction of Hung Bui, Technology Vice President at Qualcomm, highlighting the context of the discussion following Qualcomm's acquisition of VinAI Research.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-00-51.png)

**Context:** This interview on the TWiML AI Podcast features Sam Charrington speaking with Hung Bui, Technology Vice President at Qualcomm, who recently joined Qualcomm following the acquisition of VinAI Research. The discussion centers on Qualcomm's strategy for developing high-efficiency diffusion models capable of running generative AI tasks directly on mobile devices, contrasting this approach with the reliance on large cloud-based models.

## Detailed Analysis

Hung Bui, VP of Technology at Qualcomm, discusses the strategic focus on creating highly efficient diffusion models for on-device image generation and editing, a focus spurred by Qualcomm's acquisition of VinAI Research. He notes that VinAI Research was highly ranked and that the team's work led to significant achievements, such as creating a 7-billion parameter model that outperforms a 70-billion parameter model it was distilled from. The core technical approach involves knowledge distillation, where a large teacher model trains a smaller student model to retain high performance while drastically reducing size and computational requirements. This efficiency is crucial because it allows these powerful models to run locally on mobile devices, addressing concerns around inference latency and data privacy, as sensitive data does not need to leave the device. Bui mentions that this approach allows for faster iteration cycles and better personalization compared to relying solely on cloud-based large language models (LLMs) like GPT-3. He emphasizes that the goal is to achieve the same or better quality outputs with far fewer parameters and faster inference times, making advanced AI accessible everywhere.

### Guest Background and Acquisition

- Hung Bui recently joined Qualcomm following the acquisition of VinAI Research, which was ranked among the top 25 industrial AI labs based on research output
- Bui previously worked at places like Google DeepMind and Adobe Research, focusing on topics like machine learning and probabilistic reasoning
- The Vietnam AI Research Center, formerly part of SRI International, has been active since 2019.

### Efficient Model Development Strategy

- The team focused on distilling knowledge from very large models (like 70B parameters) into much smaller, efficient models (like 7B parameters) using techniques like distillation
- This approach is necessary because deploying massive models on-device is computationally prohibitive and introduces latency.

### Image Generation and Efficiency

- The smaller models achieve performance comparable to, or even better than, their much larger counterparts on specific tasks like image generation (e.g., using a distilled diffusion model)
- This efficiency is key for enabling on-device AI capabilities without heavy reliance on the cloud.

### Privacy and Latency Benefits

- Running models on-device inherently improves privacy by keeping user data local, and reduces inference latency, which is critical for real-time applications like on-device image editing.

### Future Vision and Goals

- The goal is to continue iterating on this efficient architecture to enable sophisticated AI tasks directly on mobile devices, moving beyond just text-based LLMs to include multimodal capabilities.

![Screenshot at 00:00: Introduction of Hung Bui on the TWiML AI Podcast.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-00-00.png)
![Screenshot at 00:51: Introduction slide naming Hung Bui as Technology Vice President at Qualcomm, following the acquisition of VinAI Research.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-00-51.png)
![Screenshot at 01:13: Host Sam Charrington welcomes Hung Bui to the podcast.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-01-13.png)
![Screenshot at 02:22: Split screen view of both speakers during the discussion.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-02-22.png)
![Screenshot at 04:44: Sam Charrington asks Hung Bui about his background and entry into AI.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-04-44.png)
![Screenshot at 09:57: Hung Bui gestures while explaining the challenges of large model deployment.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-09-57.png)
![Screenshot at 13:33: Hung Bui illustrates the concept of running a model \(like a 7B parameter model\) that performs as well as a much larger one.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-13-33.png)
![Screenshot at 27:27: Sam Charrington discusses the trade-off between model size and inference efficiency.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-27-27.png)
![Screenshot at 33:45: Hung Bui demonstrates the concept of applying the inverted network \(the decoder\) to go from noise back to the image.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-33-45.png)
![Screenshot at 40:44: Hung Bui discusses the importance of efficiency and data/computational constraints for on-device AI deployment.](https://ss.rapidrecap.app/screens/klk9oher8Bo/00-40-44.png)
