# MBZUAI K2-V2: A 360-Open, Reasoning-Enhanced LLM

Source: https://www.youtube.com/watch?v=SugHYYqU47E
Recap page: https://rapidrecap.app/video/SugHYYqU47E
Generated: 2026-01-28T16:37:11.215+00:00

---
## Quick Overview

The MBZUAI K2-V2 model, announced on Tuesday, September 26, 2023, is a 360-billion parameter Large Language Model that deviates from the norm by being completely open-source, including its training data logs and checkpoints, and excels particularly in reasoning tasks due to its three-phase training approach.

**Key Points:**
- K2-V2 is a new 360-billion parameter Large Language Model released by Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) on Tuesday, September 26, 2023.
- The model is completely open-source, providing full transparency by releasing training data logs, checkpoints, and the exact data recipes used.
- It outperforms both Llama 3.1 70B and Qwen 1.5 112B on reasoning tasks, achieving 92.9% accuracy on 4-person puzzles.
- The training utilized a specific three-phase approach: pre-training on a large dataset, mid-training with a new dataset, and supervised fine-tuning (SFT).
- K2-V2 uses a decaying learning rate schedule that stops before hitting zero, which the researchers found helped prevent training stagnation.
- The model's performance on reasoning tasks suggests that, unlike larger general models, it is specifically optimized for logical deduction rather than broad general knowledge.

![Screenshot at 00:20: The introduction of K2-V2, a new 70-billion parameter large language model, highlighting the model's size and the open-source nature of its release.](https://ss.rapidrecap.app/screens/SugHYYqU47E/00-00-20.jpg)

**Context:** The video discusses the release of K2-V2, a significant new large language model developed by researchers at Mohamed bin Zayed University of Artificial Intelligence (MBZUAI). Unlike many proprietary models, K2-V2 is noteworthy for its commitment to full transparency and open-sourcing, setting it apart from models like Llama and Qwen, especially in its approach to reasoning capabilities.

## Detailed Analysis

The MBZUAI K2-V2 model, a 360-billion parameter LLM, was released on September 26, 2023, with a focus on complete openness, providing the entire blueprint including training data logs and checkpoints, which contrasts with proprietary models that keep their methods secret. The researchers argue this transparency bridges the gap between general-purpose models and specialized reasoning models. The model was trained using a three-phase approach: pre-training, mid-training, and supervised fine-tuning (SFT). Crucially, they employed a decaying learning rate schedule that stopped just before reaching zero, preventing the model from overfitting or stagnating, which they compare to setting a car's transmission to a usable gear instead of letting it coast to zero. This tailored approach led to superior performance on reasoning benchmarks; K2-V2 scored 92.9% accuracy on 4-person puzzles, outperforming Llama 3.1 70B and Qwen 1.5 112B. Furthermore, the model demonstrated an improved context window handling, allowing it to process large documents without breaking its internal order, unlike models that require external tools like calculators for complex math. The high scores on reasoning tasks suggest that K2-V2 is intentionally optimized as a problem-solver rather than a general knowledge encyclopedia, which is reflected in its performance metrics across different reasoning benchmarks (e.g., MMLU, GSM8K). The model's success in reasoning suggests that simply increasing model size (the 'bigger is better' philosophy) is not the only path to improvement, as K2-V2 achieved strong results without the massive computational cost of models like GPT-4, whose word count is far higher.

### K2-V2 Model Overview

- Released 09/26/2023
- 360 billion parameters
- Completely open-source blueprint, logs, and checkpoints
- Explicitly designed as a specialized reasoning model, not a general knowledge engine

### Training Methodology

- Three phases used: pre-training (TXT360 dataset), mid-training (new dataset), and supervised fine-tuning (SFT)
- Used a decaying learning rate schedule that stopped before zero to maintain plasticity

### Performance Benchmarks

- Scored 92.9% accuracy on 4-person puzzles, surpassing Llama 3.1 70B and Qwen 1.5 112B
- Scored high on reasoning tasks (e.g., MMLU, GSM8K)
- High-effort setting performed worse than medium setting on some tasks, suggesting it overthinks

### Transparency and Openness

- Full blueprint released, allowing others to verify results and build upon the work
- Open sourcing counters the trend of proprietary models hiding their training recipes and data sources

### Reasoning vs. General Knowledge

- Model excels at logical deduction and tasks requiring multi-step reasoning (like function calling)
- Does not aim to be a general encyclopedia; struggles with simple factual recall (e.g., capitals)

![Screenshot at 00:00: Video intro screen featuring two podcasters and the text 'BECOME A MEMBER TODAY!' over a grid background.](https://ss.rapidrecap.app/screens/SugHYYqU47E/00-00-00.jpg)
![Screenshot at 00:15: Visual confirmation of the model name and size: 'MBZUAI K2-V2' and '70 billion parameter large language model'.](https://ss.rapidrecap.app/screens/SugHYYqU47E/00-00-15.jpg)
![Screenshot at 01:35: A slide or graphic summarizing the three-phase training approach: Pre-training, Mid-training, and Supervised Fine-Tuning \(SFT\).](https://ss.rapidrecap.app/screens/SugHYYqU47E/00-01-35.jpg)
![Screenshot at 05:56: A comparison point noting that while high-effort queries yield better answers, the paper points out a 'catch-22' where larger context windows can sometimes lead to poorer performance.](https://ss.rapidrecap.app/screens/SugHYYqU47E/00-05-56.jpg)
![Screenshot at 07:47: A graphic showing the performance comparison on reasoning tasks, stating K2-V2 achieved 92.9% accuracy on 4-person puzzles, outperforming Llama and Qwen models.](https://ss.rapidrecap.app/screens/SugHYYqU47E/00-07-47.jpg)
