NVIDIA's New Open Models

Quick Overview

NVIDIA announced its next-generation AI platform, the Vera Rubin GPU, delivering significant performance boosts over the Blackwell generation (e.g., 5x inference, 3.5x training), alongside new open models like Nemotron for speech/multimodal AI and Alpamao for autonomous vehicles, all while emphasizing falling inference costs and broad ecosystem adoption.

Key Points: The NVIDIA Vera Rubin GPU offers major generational leaps over Blackwell, including 50 PFLOPS for NVP4 Inference (5x Blackwell), 35 PFLOPS for NVP4 Training (3.5x Blackwell), and 336 billion transistors (1.6x Blackwell). NVIDIA announced 13 new open models at CES 2026, including the NVIDIA Alpamao, an Open Reasoning VLA for autonomous vehicles, which integrates egomotion, multi-camera video, and user commands to output driving decisions and trajectory. The Nemotron family of open models was released, featuring Nemotron Speech (delivering 10x faster performance than other models in its class for real-time ASR), Nemotron RAG for multimodal retrieval-augmented generation, and Nemotron Safety models. The presentation highlighted the trend of falling token costs (10x cheaper per year) and rapidly scaling model sizes (10x parameters per year), necessitating more powerful hardware like the announced chips. NVIDIA is shipping its Full-Stack AV solution on the 2025 Mercedes-Benz CLA, featuring the Alpamao reasoning model on top of a classical AV stack and safety OS. New open models for Physical AI (Cosmos), robotics (Isaac GROOT), and biomedical applications (Clara) were introduced, supported by vast open datasets, including 10 trillion language tokens and 500,000 robotics trajectories. The company showcased the cost-effectiveness of running the streaming Nemotron Speech model locally (on a Mac) with low latency due to cache-aware streaming ASR architecture.

Context: The video captures a keynote presentation, likely from NVIDIA's annual event (referenced as CES 2026), where CEO Jensen Huang unveils major advancements across their AI and computing stack. The presentation focuses heavily on the next generation of hardware (Rubin GPU), new foundational open-source models for various domains (autonomous vehicles, speech, robotics, healthcare), and the continued trend of increasing AI demands necessitating superior compute power.

Raw markdown version of this recap