# To The “Muon”? A Slightly Technical Breakdown of Kimi K2

Source: https://www.youtube.com/watch?v=m1PUl-aUk8c
Recap page: https://rapidrecap.app/video/m1PUl-aUk8c
Generated: 2025-08-26T15:33:22.03+00:00

---
## Quick Overview

Moonshot AI's Kimi K2 language model, trained with the Muon optimizer, achieves state-of-the-art performance on open-source non-thinking models, excelling in agentic capabilities and various benchmarks like Tau2-Bench and LiveCodeBench. The model's training cost is estimated at $20-30 million, utilizing 15.5 trillion tokens and 32.6 billion activated parameters.

**Key Points:**
- Kimi K2, a Mixture-of-Experts (MoE) model, achieves state-of-the-art performance among open-source non-thinking models, with strengths in agentic capabilities.
- It scores highly on benchmarks like Tau2-Bench (66.1), ACEBench (En) (76.5), SWE-Bench Verified (65.8), SWE-Bench Multilingual (47.3), LiveCodeBench v6 (53.7), AIME 2025 (49.5), GPQA-Diamond (75.1), and OJBench (27.1).
- The model was trained using the Muon optimizer, which improves training stability and efficiency.
- The training process involved 15.5 trillion tokens and resulted in a cost estimate of $20-30 million, utilizing 32.6 billion activated parameters.
- Kimi K2 is open-sourced, with a base model and an instruct-tuned version available.
- The Muon optimizer itself has shown significant improvements, setting new speed records in tasks like NanoGPT speedrunning, improving training speed by 35% compared to AdamW.
- The paper introducing Muon was released on February 24, 2025, detailing its scalability and effectiveness.

![Screenshot at 00:00: The title card introduces Moonshot AI and the concept of the 'Long LLM era', setting the stage for the discussion of Kimi K2.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-00-00.png)

**Context:** This video breaks down the technical aspects and performance of Kimi K2, a large language model developed by Moonshot AI. It highlights the model's architecture, its training process using the Muon optimizer, and its benchmark results, comparing it to other leading models. The discussion also touches upon the cost of training such a model and the underlying techniques that enable its performance.

## Detailed Analysis

Moonshot AI's Kimi K2 is a 1.04 trillion-parameter Mixture-of-Experts (MoE) transformer model with 32 billion activated parameters, designed to excel in agentic capabilities. Trained using the novel Muon optimizer, Kimi K2 achieves state-of-the-art performance across various benchmarks, including Tau2-Bench, ACEBench, SWE-Bench Verified, SWE-Bench Multilingual, LiveCodeBench v6, AIME 2025, GPQA-Diamond, and OJBench. The Muon optimizer, introduced in a paper released on February 24, 2025, has demonstrated significant improvements in training speed and stability, setting new records in tasks like NanoGPT speedrunning by enhancing training speed by 35% over AdamW. The training process for Kimi K2 utilized 15.5 trillion tokens, with an estimated cost of $20-30 million. The model's architecture, which includes 384 experts and a reduced number of attention heads (64 compared to 128 in DeepSeek-V3), contributes to its efficiency and performance. The research highlights that increased sparsity in MoE models consistently lowers training and validation loss, enhancing overall model performance. Kimi K2 is released as open-source, providing researchers and developers with access to its base and instruct-tuned checkpoints.

### Model Introduction

- Kimi K2 is a 1.04T parameter MoE LLM with 32B activated parameters, designed for agentic capabilities
- Its architecture is similar to DeepSeek-V3 but with more experts (384 vs 256) and fewer attention heads (64 vs 128)
- Scaling law analysis reveals increased sparsity yields substantial performance improvements

### Optimizer

- Muon optimizer improves training stability and efficiency
- Achieved new speed records in NanoGPT speedrunning, improving speed by 35% over AdamW
- Muon's training loss graph shows a stable descent without significant spikes

### Performance Benchmarks

- Kimi K2 achieves state-of-the-art performance on open-source non-thinking models
- Scores high on Tau2-Bench (66.1), ACEBench (76.5), SWE-Bench Verified (65.8), SWE-Bench Multilingual (47.3), LiveCodeBench v6 (53.7), AIME 2025 (49.5), GPQA-Diamond (75.1), OJBench (27.1)

### Training Cost

- Estimated at $20-30 million
- Utilized 15.5T tokens and 32.6B activated parameters

### Open Source

- Kimi K2 is open-sourced with base and instruct-tuned model checkpoints available

### Sparsity Scaling Law

- Increasing sparsity (total experts) leads to lower training/validation loss and better performance
- Kimi K2 uses a sparsity of 48 (8 out of 384 experts per forward pass) for optimal balance

### Architectural Comparison

- Kimi K2 vs DeepSeek-V3
- Kimi K2 has 54% more parameters (1.04T vs 671B), 13% fewer activated parameters (32.6B vs 37B), 50% more experts (384 vs 256), 50% fewer attention heads (64 vs 128), and 67% fewer dense layers (1 vs 3)
- Kimi K2 does not use expert grouping, allowing routers to spread across GPUs

![Screenshot at 00:00: The title card introduces Moonshot AI and the concept of the 'Long LLM era', setting the stage for the discussion of Kimi K2.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-00-00.png)
![Screenshot at 00:05: The introduction slide for Kimi K2, highlighting its name and 'Open Agentic Intelligence'.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-00-05.png)
![Screenshot at 00:20: A bar chart comparing Kimi K2's performance against other models on various benchmarks like GPQA, AIME25, and LiveCodeBench v6.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-00-20.png)
![Screenshot at 00:34: A screenshot of the EQ-Bench 3 leaderboard showing Kimi K2's performance in emotional intelligence benchmarks.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-00-34.png)
![Screenshot at 00:40: A leaderboard ranking open models by provider \(Text\), with Kimi K2 listed at the top with an Arena Score of 1420.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-00-40.png)
![Screenshot at 00:42: A bar chart showing GPQA Diamond Benchmark Leaderboard results, with Kimi K2 scoring 76.6%.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-00-42.png)
![Screenshot at 00:51: A comprehensive performance comparison across multiple benchmarks, including Agentic and Competitive Coding, Tool Use, and Math & STEM.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-00-51.png)
![Screenshot at 01:18: A breakdown of the estimated training cost for Kimi K2, showing a total estimate of $20-30M.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-01-18.png)
![Screenshot at 01:25: Text highlighting Kimi K2's state-of-the-art performance among open-source non-thinking models, particularly in agentic capabilities.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-01-25.png)
![Screenshot at 02:57: A loss vs. tokens graph illustrating the training progress of Kimi K2, showing a steady decrease in loss over billions of tokens, attributed to the Muon optimizer.](https://ss.rapidrecap.app/screens/m1PUl-aUk8c/00-02-57.png)
