# Foundation models for autonomous driving

Source: https://www.youtube.com/watch?v=w53P2_LozEI
Recap page: https://rapidrecap.app/video/w53P2_LozEI
Generated: 2026-02-13T15:09:20.405+00:00

---
## Quick Overview

Waymo's autonomous driving technology achieves superhuman sensing ability by fusing multimodal data (vision, Lidar, Radar) and employing Large Multimodal Models (LMMs) based on Gemini architecture to accurately predict complex scenarios and align with real-world human safety preferences, resulting in significant reductions in crashes involving vulnerable road users compared to human drivers.

**Key Points:**
- Waymo's system achieves superhuman sensing ability by integrating vision, Lidar, and Radar data using a Large Multimodal Model based on Gemini.
- The model treats driving as a conversation, using trajectories as sentences and motion points as vocabulary, leveraging local continuity and global context.
- Post-training alignment using expert preferences (Inverse Reinforcement Learning) corrects pre-trained model behaviors, such as yielding to pedestrians and adopting human-like stopping behavior.
- Compared to human drivers over 56.3 million miles, Waymo Driver had 92% fewer crashes with injuries to pedestrians, 82% fewer with injuries to cyclists, and 82% fewer with injuries to motorcyclists.
- The system reduces the frequency of property damage claims by 88% and bodily injury claims by 92% based on third-party evaluation.
- Waymo is actively expanding operations, planning to launch in Miami next year and scanning many other cities for future deployment, including Japan.
- The approach focuses on aligning the AI with real-world safety objectives (safety, comfort, compliance) rather than simply imitating human driving.

![Screenshot at 00:04: The opening graphic for the AI for Good Global Summit, scheduled for 8-11 July 2025 in Geneva, Switzerland, establishes the context of the presentation focusing on AI applications for global challenges.](https://ss.rapidrecap.app/screens/w53P2_LozEI/00-00-04.jpg)

**Context:** The presentation focuses on Waymo's advancements in autonomous driving AI, specifically detailing how they achieve safety and reliability, particularly when dealing with complex, long-tail scenarios. The speaker, an engineer from Waymo, introduces the use of multimodal AI models, specifically those based on the Gemini architecture, to process sensor data (Lidar, vision, radar) and textual context (like high-level commands or road sign interpretations) to generate accurate trajectories and ensure safety alignment with human preferences.

## Detailed Analysis

The speaker, an engineer from Waymo, details the development of their autonomous driving system, emphasizing the goal to be the world's most trusted driver. This trust is built upon achieving safety, consistency, and predictability, which is critical given the long tail of complex scenarios they must handle, such as unusual behavior, foreign objects, extreme weather, and unique interactions (like interacting with police). Waymo's approach involves a Large Multimodal Model based on Gemini, which fuses data from Lidar, vision (cameras), and text (high-level commands, historical status) to predict trajectories. The system utilizes scaling laws derived from motion forecasting research, showing predictable performance gains with increased model size and data. A key technique for safety alignment involves using Inverse Reinforcement Learning, where experts rank model responses (like yielding to pedestrians or stopping behavior) to fine-tune the model's objectives beyond simply imitating humans. This results in significantly better safety metrics: 92% fewer crashes with pedestrian injuries, 82% fewer with cyclist injuries, and 82% fewer with motorcyclist injuries compared to human drivers over 56.3 million miles. Furthermore, third-party evaluation showed an 88% reduction in property damage claims and a 92% reduction in bodily injury claims. The system also demonstrates advanced language understanding by correctly interpreting complex parking signs based on current context (day/time). Waymo is actively expanding its operational footprint, with plans for Miami next year and continued scouting for new markets, including Japan.

### Waymo's Mission and Scale

- Aim to be the world's most trusted driver
- Currently serving over 250,000 paid trips per week
- Expanding operations to Miami and scouting other cities, including Japan.

### Safety Performance Metrics (Compared to human drivers over 56.3M miles)

- 92% fewer crashes with pedestrian injuries
- 82% fewer crashes with cyclist injuries
- 82% fewer crashes with motorcyclist injuries
- 88% reduction in property damage claims.

### Driving as a Conversation (Motion Forecasting)

- Trajectories treated as sentences in a new language
- Vocabulary built from motion points (vectors)
- Model architecture similar to a Large Language Model (LLM).

### Model Alignment via Expert Preferences

- Uses Inverse Reinforcement Learning to infer true objectives (safety, comfort, compliance) from expert rankings of MotionLM responses
- This prevents models trained only on imitation from exhibiting risky behaviors (e.g., stopping too late or driving too close to pedestrians).

### Superhuman Sensing Ability

- Demonstrated by Lidar successfully perceiving objects (pedestrians) in extremely low visibility conditions (heavy fog) where HD cameras fail.

### What Makes Autonomous-Vehicle AI Challenging

- Complex physical environment
- High-performance requirements
- Real-time computation.

### Multimodal Models for Motion Prediction (EMMA)

- Router directs high-level commands and contextual text status
- Fuses text and vision data via a Large Multimodal Model based on Gemini to output Trajectories.

![Screenshot at 00:02: The initial colorful, abstract animation resolves into the AI for Good Global Summit logo, signaling the start of the event presentation.](https://ss.rapidrecap.app/screens/w53P2_LozEI/00-00-02.jpg)
![Screenshot at 00:24: Waymo's mission statement, "Be the world's most trusted driver," displayed above a Waymo autonomous vehicle, setting the core theme of the talk.](https://ss.rapidrecap.app/screens/w53P2_LozEI/00-00-24.jpg)
![Screenshot at 01:15: A graphic showcasing Waymo's current operational service areas \(SF, PHX, LA, ATX, ATL\) alongside their respective service maps, detailing expansion.](https://ss.rapidrecap.app/screens/w53P2_LozEI/00-01-15.jpg)
![Screenshot at 02:07: A slide highlighting Waymo One's scale, stating they are "Now serving over 250,000 paid trips per week," illustrating the volume of real-world testing.](https://ss.rapidrecap.app/screens/w53P2_LozEI/00-02-07.jpg)
![Screenshot at 07:26: A comparison showing how Waymo's system handles complex scenarios like yielding to pedestrians and human-like stopping behavior in simulations, contrasting pre-trained model failures \(red circles\) with post-training alignment success \(green circles\).](https://ss.rapidrecap.app/screens/w53P2_LozEI/00-07-26.jpg)
