NVIDIA's NEW Nemotron 3 Super in 6 Minutes

Quick Overview

NVIDIA's Nemotron 3 Super, a 120B-parameter model with 12B active parameters utilizing a hybrid Mamba-Transformer Mixture-of-Experts (MoE) architecture, delivers superior performance, notably achieving 5x higher throughput than previous Nemotron models and leading intelligence/efficiency benchmarks like the Artificial Analysis Openness Index (scoring 83) and the Intelligence vs. Efficiency chart.

Key Points: Nemotron 3 Super is a 120-billion-parameter open model featuring 12 billion active parameters via a hybrid Mamba-Transformer MoE architecture. The model achieved a score of 36 on the Artificial Analysis Intelligence Index and an 83 on the Artificial Analysis Openness Index, placing it in the most attractive quadrant for openness and intelligence. In throughput benchmarks (00:00), Nemotron 3 Super (BF16 configuration) shows significant throughput gains, reaching 2.2 tokens/s/GPU on the ISV/IDSL benchmark, compared to lower values for peers. The architecture uses a 'Latent Twist' MoE approach where tokens are compressed before routing, allowing for 4x more experts at the same cost compared to standard MoE. Nemotron 3 Super boasts a 1-million-token context window, enabling agents to retain full workflow state in memory and prevent goal drift (4:04). The model is trained on synthetic data generated using frontier reasoning models and is released under a permissive license, allowing deployment on workstations, data centers, or cloud infrastructure (4:14).

Context: The video discusses the release and technical details of NVIDIA's new large language model, Nemotron 3 Super. This model is positioned as a significant step forward, combining a Transformer architecture with Mamba components in a Mixture-of-Experts (MoE) setup, specifically employing a 'Latent MoE' twist for efficiency. The presentation relies heavily on benchmark charts from Artificial Analysis comparing Nemotron 3 Super against various open-source peers across metrics like intelligence, openness, and inference throughput, highlighting its competitive positioning.

Raw markdown version of this recap