# NVIDIA's New Open Models

Source: https://www.youtube.com/watch?v=JrCKb6gM9KM
Recap page: https://rapidrecap.app/video/JrCKb6gM9KM
Generated: 2026-01-06T15:07:29.334+00:00

---
## Quick Overview

NVIDIA announced its next-generation AI platform, the Vera Rubin GPU, delivering significant performance boosts over the Blackwell generation (e.g., 5x inference, 3.5x training), alongside new open models like Nemotron for speech/multimodal AI and Alpamao for autonomous vehicles, all while emphasizing falling inference costs and broad ecosystem adoption.

**Key Points:**
- The NVIDIA Vera Rubin GPU offers major generational leaps over Blackwell, including 50 PFLOPS for NVP4 Inference (5x Blackwell), 35 PFLOPS for NVP4 Training (3.5x Blackwell), and 336 billion transistors (1.6x Blackwell).
- NVIDIA announced 13 new open models at CES 2026, including the NVIDIA Alpamao, an Open Reasoning VLA for autonomous vehicles, which integrates egomotion, multi-camera video, and user commands to output driving decisions and trajectory.
- The Nemotron family of open models was released, featuring Nemotron Speech (delivering 10x faster performance than other models in its class for real-time ASR), Nemotron RAG for multimodal retrieval-augmented generation, and Nemotron Safety models.
- The presentation highlighted the trend of falling token costs (10x cheaper per year) and rapidly scaling model sizes (10x parameters per year), necessitating more powerful hardware like the announced chips.
- NVIDIA is shipping its Full-Stack AV solution on the 2025 Mercedes-Benz CLA, featuring the Alpamao reasoning model on top of a classical AV stack and safety OS.
- New open models for Physical AI (Cosmos), robotics (Isaac GROOT), and biomedical applications (Clara) were introduced, supported by vast open datasets, including 10 trillion language tokens and 500,000 robotics trajectories.
- The company showcased the cost-effectiveness of running the streaming Nemotron Speech model locally (on a Mac) with low latency due to cache-aware streaming ASR architecture.

![Screenshot at 00:09: A diagram illustrating the 'NVIDIA Full-Stack Physical AI Platform' showing the flow from GB300 training hardware through Cosmos/Omniverse to inference \(THOR\) and simulation \(RTX PRO\), contextualizing the broader ecosystem for the announced models.](https://ss.rapidrecap.app/screens/JrCKb6gM9KM/00-00-09.jpg)

**Context:** The video captures a keynote presentation, likely from NVIDIA's annual event (referenced as CES 2026), where CEO Jensen Huang unveils major advancements across their AI and computing stack. The presentation focuses heavily on the next generation of hardware (Rubin GPU), new foundational open-source models for various domains (autonomous vehicles, speech, robotics, healthcare), and the continued trend of increasing AI demands necessitating superior compute power.

## Detailed Analysis

Jensen Huang opened the presentation by welcoming everyone to the Consumer Electronics Show (CES) 2026, immediately transitioning to the NVIDIA Full-Stack Physical AI Platform diagram, which connects training (GB300), simulation (Omniverse/RTX Pro), inference (THOR), and specialized models like Cosmos, Alpamao, and GROOT. He discussed the insane demand for AI computing, showing charts that illustrate model size growing 10x per year, test-time thinking scaling 5x per year, while token costs drop 10x cheaper per year. The first major hardware announcement was the Vera Rubin GPU, which is 5x faster for inference and 3.5x faster for training compared to the Blackwell generation, featuring 336 billion transistors. Huang highlighted that major hyperscalers are already asking for these chips for data centers. He then detailed the new open models, starting with NVIDIA Alpamao, an open reasoning Vision Language Agent (VLA) for autonomous vehicles, which processes egomotion, video, and user commands to produce driving decisions and trajectory predictions, demonstrated in simulation. He also announced that NVIDIA ships its full-stack AV solution on the 2025 Mercedes-Benz CLA. Next, he detailed the ecosystem support, referencing quotes from leaders at AWS, OpenAI, Meta, Microsoft, Google, and Oracle praising the Rubin platform. For robotics, he mentioned Isaac GROOT models trained via Cosmos synthetic data, showing a transfer from simulation to a real-world robot task. For healthcare, the Clara AI models were introduced, including La-Proteina, ReaSyn v2, KERMT, and RNAPro, aimed at accelerating drug discovery with a dataset of 455,000 synthetic protein structures. Finally, he detailed the Nemotron family of open models, spanning Speech, RAG, and Safety. The Nemotron Speech model achieves leaderboard-topping performance, delivering 10x faster performance than others in its class for low-latency speech recognition, which can be run locally with low latency due to cache-aware streaming architecture.

### CES 2026 Kickoff

- Welcome to the show
- Introduction of the Full-Stack Physical AI Platform
- Discussion of scaling trends (model size, token cost)

### Hardware Announcement

- Introduction of the NVIDIA Vera Rubin GPU
- 5x inference speed and 3.5x training speed improvement over Blackwell
- 336 Billion transistors

### Autonomous Vehicles & Robotics

- Announcement of NVIDIA Alpamao (Open Reasoning VLA) for AVs
- Full-Stack AV shipping on 2025 Mercedes-Benz CLA
- Demonstration of Isaac GROOT training via Cosmos simulation

### AI Ecosystem Support

- Quotes from CEOs of OpenAI, Google, Meta, Microsoft, and Oracle endorsing the Rubin platform
- Mention of 13 open models released

### Healthcare & Life Sciences

- Launch of new Clara AI models (La-Proteina, ReaSyn v2, KERMT, RNAPro) for drug discovery
- Use of 455,000 synthetic protein structures dataset

### Nemotron Models

- Release of Nemotron family for speech, RAG, and safety
- Nemotron Speech offers 10x faster performance in ASR via cache-aware streaming
- Models support multimodal data and local execution

![Screenshot at 00:00: Establishing shot of the venue at night, showing the large screen displaying a futuristic building rendering.](https://ss.rapidrecap.app/screens/JrCKb6gM9KM/00-00-00.jpg)
![Screenshot at 00:10: Diagram detailing the 'NVIDIA Full-Stack Physical AI Platform' pipeline, illustrating the connection between training \(GB300\), simulation \(Omniverse\), and output branches \(Inference/THOR, Simulation/RTX PRO\).](https://ss.rapidrecap.app/screens/JrCKb6gM9KM/00-00-10.jpg)
![Screenshot at 00:45: Slide announcing 'Vera Rubin: Six New Chips – One Giant Leap to the Next Frontier' displayed above a massive server rack visualization.](https://ss.rapidrecap.app/screens/JrCKb6gM9KM/00-00-45.jpg)
![Screenshot at 01:23: Three charts summarizing the 'Insane Demand for AI Computing,' showing Model Size Growing \(10x/year\), Test-Time Scaling \(5x/year\), and Token Cost \(10x Cheaper/Year\).](https://ss.rapidrecap.app/screens/JrCKb6gM9KM/00-01-23.jpg)
![Screenshot at 03:30: In-car view demonstrating the NVIDIA Alpamao system in a Mercedes-Benz, showing trajectory prediction and reasoning \('Unprotected Left, Yield to Pedestrians and Traffic'\).](https://ss.rapidrecap.app/screens/JrCKb6gM9KM/00-03-30.jpg)
