# NVIDIA's NEW Nemotron 3 Super in 6 Minutes

Source: https://www.youtube.com/watch?v=JNAvKGU2mOo
Recap page: https://rapidrecap.app/video/JNAvKGU2mOo
Generated: 2026-03-12T13:34:14.54+00:00

---
## Quick Overview

NVIDIA's Nemotron 3 Super, a 120B-parameter model with 12B active parameters utilizing a hybrid Mamba-Transformer Mixture-of-Experts (MoE) architecture, delivers superior performance, notably achieving 5x higher throughput than previous Nemotron models and leading intelligence/efficiency benchmarks like the Artificial Analysis Openness Index (scoring 83) and the Intelligence vs. Efficiency chart.

**Key Points:**
- Nemotron 3 Super is a 120-billion-parameter open model featuring 12 billion active parameters via a hybrid Mamba-Transformer MoE architecture.
- The model achieved a score of 36 on the Artificial Analysis Intelligence Index and an 83 on the Artificial Analysis Openness Index, placing it in the most attractive quadrant for openness and intelligence.
- In throughput benchmarks (00:00), Nemotron 3 Super (BF16 configuration) shows significant throughput gains, reaching 2.2 tokens/s/GPU on the ISV/IDSL benchmark, compared to lower values for peers.
- The architecture uses a 'Latent Twist' MoE approach where tokens are compressed before routing, allowing for 4x more experts at the same cost compared to standard MoE.
- Nemotron 3 Super boasts a 1-million-token context window, enabling agents to retain full workflow state in memory and prevent goal drift (4:04).
- The model is trained on synthetic data generated using frontier reasoning models and is released under a permissive license, allowing deployment on workstations, data centers, or cloud infrastructure (4:14).

![Screenshot at 0:00: A bar chart comparing the Accuracy and Throughput of Nemotron 3 Super \(BF16 and NVFP4\), GPT-OSS-120B, and Qwen 3.5-122B across various benchmarks, showing Nemotron 3 Super's relative performance.](https://ss.rapidrecap.app/screens/JNAvKGU2mOo/00-00-00.jpg)

**Context:** The video discusses the release and technical details of NVIDIA's new large language model, Nemotron 3 Super. This model is positioned as a significant step forward, combining a Transformer architecture with Mamba components in a Mixture-of-Experts (MoE) setup, specifically employing a 'Latent MoE' twist for efficiency. The presentation relies heavily on benchmark charts from Artificial Analysis comparing Nemotron 3 Super against various open-source peers across metrics like intelligence, openness, and inference throughput, highlighting its competitive positioning.

## Detailed Analysis

The video introduces NVIDIA's Nemotron 3 Super, a 120-billion-parameter model with only 12 billion active parameters, utilizing a novel hybrid Mamba-Transformer Mixture of Experts (MoE) architecture. The presenter explains that this architecture employs a 'Latent Twist,' where tokens are compressed before routing, allowing for 4 times more experts at the same computational cost compared to standard MoE models (0:51). This efficiency allows the model to perform well while maintaining high intelligence, as evidenced by its top placement in the 'most attractive quadrant' on the Artificial Analysis Openness vs. Intelligence Index (1:55). The model also features a massive 1-million-token context window, crucial for complex, multi-agent workflows where retaining full state history prevents goal drift (4:04). Benchmark results shown, such as on the Intelligence vs. Efficiency ratio (3:04), confirm that Nemotron 3 Super achieves superior throughput (e.g., 11% higher than gpt-oss-120b on a specific test) while maintaining high intelligence scores across coding, math, and science evaluations (2:11). The model is open-sourced under a permissive license, and NVIDIA provides the full methodology, synthetic training data (over 10 trillion tokens), and 15 training environments for further fine-tuning via the NeMo platform (4:14). The model is accessible via various cloud providers like Google Cloud Vertex AI, Oracle Cloud Infrastructure, AWS, and Azure (4:45).

### Performance Benchmarks (Accuracy/Throughput)

- Nemotron 3 Super (BF16) achieved high accuracy across benchmarks like HMMF Feb'25 Math (94.9%) and showed superior throughput, reaching 2.2 tokens/s/GPU on ISV/IDSL (0:00).

### Mixture of Experts (MoE) Architecture

- Introduced as a 'Latent Twist' MoE, which compresses tokens before routing, enabling 4x more experts at the same cost compared to Standard MoE (0:34).

### Openness and Intelligence Ranking

- Nemotron 3 Super scored 83 on the Artificial Analysis Openness Index, placing it significantly above peers in the 'Most Attractive Quadrant' for Openness and Intelligence (1:55).

### Context Window and Agentic Capability

- Features a 1-million-token context window, allowing agents to retain full workflow state and prevent goal drift during long tasks (4:04).

### Accessibility and Training Data

- Released under a permissive license, with complete methodology and over 10 trillion tokens of synthetic pre- and post-training data available (4:14).

### Efficiency Testing

- Self-hosted throughput tests showed Nemotron 3 Super (NVFP4) achieved 11% higher throughput per NVIDIA B200 GPU than gpt-oss-120b (MXFP4) in a specific test scenario (5:24).

![Screenshot at 0:00: A bar chart comparing the Accuracy and Throughput of Nemotron 3 Super \(BF16 and NVFP4\), GPT-OSS-120B, and Qwen 3.5-122B across various benchmarks, showing Nemotron 3 Super's relative performance.](https://ss.rapidrecap.app/screens/JNAvKGU2mOo/00-00-00.jpg)
![Screenshot at 0:17: A diagram explaining the Mixture of Experts \(MoE\) architecture, contrasting standard routing of raw tokens with the MoE router picking a small subset of specialists for an input token.](https://ss.rapidrecap.app/screens/JNAvKGU2mOo/00-00-17.jpg)
![Screenshot at 0:51: A comparison slide detailing the 'Latent Twist' in MoE, where compression happens before routing, allowing for 4x more experts at the same cost compared to standard MoE.](https://ss.rapidrecap.app/screens/JNAvKGU2mOo/00-00-51.jpg)
![Screenshot at 2:16: A scatter plot ranking models based on Artificial Analysis Openness Index \(Y-axis\) versus Artificial Analysis Intelligence Index \(X-axis\), highlighting the 'Most Attractive Quadrant' for high performance and openness.](https://ss.rapidrecap.app/screens/JNAvKGU2mOo/00-02-16.jpg)
