# Instella: Fully Open Language Models with Stellar Performance

Source: https://www.youtube.com/watch?v=pVs4XLk87oo
Recap page: https://rapidrecap.app/video/pVs4XLk87oo
Generated: 2025-11-17T15:34:24.432+00:00

---
## Quick Overview

Instella 3B achieves superior performance across benchmarks compared to similarly sized open and fully open models, notably outperforming Llama 3.5 by a significant margin on the GSM8K math reasoning benchmark (4.1% lead) and achieving the highest reported performance among all models in its size class on the GPQA benchmark, demonstrating that smaller, highly-tuned models can surpass larger, less transparent ones.

**Key Points:**
- Instella 3B outperforms Llama 3.5 by 4.1% on the GSM8K math reasoning benchmark.
- Instella 3B achieved the highest reported performance among all models in its size class on the GPQA benchmark.
- The model was trained using a two-stage process: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning from Human Feedback (RLHF) using Direct Preference Optimization (DPO).
- The training involved 128K instruction-response pairs and 58 Billion tokens of synthetic data generated using a larger teacher model.
- The training methodology focused on maximizing performance on complex reasoning tasks like Truthful QA (TQA) and General Knowledge Question Answering (GKQA).
- Instella 3B demonstrated superior performance compared to competitors like Llama 3.5, beating it by 14% on average across benchmarks.
- The model's 128K context window version achieved a 4.1% lead over Llama 3.5 on GSM8K.

![Screenshot at 03:33: The speaker discusses how Instella 3B uses an explicit technique \(DPO\) to learn to say 'no' to bad answers, validating the approach of training for helpfulness and safety over simple memorization.](https://ss.rapidrecap.app/screens/pVs4XLk87oo/00-03-33.png)

**Context:** The video introduces Instella 3B, a new large language model (LLM) developed by AMD, positioning it as a highly capable, fully open model designed to challenge existing proprietary and open-source leaders. The discussion centers on its performance metrics, particularly against models like Llama, and the innovative, multi-stage training methodology used to achieve this high level of reasoning and factual accuracy with a relatively small parameter count.

## Detailed Analysis

The Instella 3B model demonstrates stellar performance, especially for its size (3 billion parameters), setting a new standard for fully open language models. The key takeaway is that specialized training, emphasizing complex reasoning and safety alignment, allows smaller models to significantly outperform larger, less transparent competitors. AMD achieved this through a rigorous two-stage training process: Supervised Fine-Tuning (SFT) on 128K instruction-response pairs, followed by RLHF using Direct Preference Optimization (DPO) guided by human preferences. This process focused on teaching the model to reason logically, rather than just recall facts, and to refuse unsafe or nonsensical queries. Performance comparisons show Instella 3B beating Llama 3.5 by a substantial margin (e.g., 4.1% lead on GSM8K) and setting new state-of-the-art scores for its class on benchmarks like GPQA. The researchers engineered the training to focus on high-quality, diverse synthetic data generation using a larger model to create context-rich training examples, ensuring the model generalizes well to real-world tasks while maintaining a small footprint.

### Instella 3B Performance Metrics

- Outperforms Llama 3.5 on GSM8K by 4.1%
- Achieved highest reported performance for its size class on GPQA
- Beat next competitor by 14% on average across benchmarks

### Training Methodology

- Two stages: SFT on 128K instruction-response pairs followed by RLHF using DPO
- Used 58 Billion tokens of synthetic data
- Focused on teaching transferable strategic reasoning skills

### Data Curation and Scale

- Used synthetic data generated from a larger capable teacher model
- Context window extended to 128K tokens for the high-performing variant
- Explicitly avoided massive brute-force training runs

### Key Technical Advantages

- Utilizes Group Relative Policy Optimization (GRPO) for alignment
- Explicitly trained to refuse bad answers (safety alignment)
- Achieved high performance without relying on black-box proprietary data

![Screenshot at 00:01: Visual graphic showing two podcast hosts, symbolizing the AI Paper Podcast Daily format.](https://ss.rapidrecap.app/screens/pVs4XLk87oo/00-00-01.png)
![Screenshot at 00:19: Visual emphasis on the importance of the model, stating that Instella models directly challenge the idea that top-tier performance requires secrecy.](https://ss.rapidrecap.app/screens/pVs4XLk87oo/00-00-19.png)
![Screenshot at 01:13: A visual representation of the two-stage training process: one stage for the finished product, and another for the full recipe/ingredients.](https://ss.rapidrecap.app/screens/pVs4XLk87oo/00-01-13.png)
![Screenshot at 02:24: Data point showing Instella 3B's 3 Billion parameters is 'compact' compared to the competition.](https://ss.rapidrecap.app/screens/pVs4XLk87oo/00-02-24.png)
![Screenshot at 03:13: Speaker highlighting the difference between a finished product and having the entire recipe, emphasizing transparency.](https://ss.rapidrecap.app/screens/pVs4XLk87oo/00-03-13.png)
![Screenshot at 04:04: Visual representation of the data quality argument in action, contrasting the model's performance with larger, less transparent models.](https://ss.rapidrecap.app/screens/pVs4XLk87oo/00-04-04.png)
![Screenshot at 05:55: Discussion about smoothing out noise and boosting robustness without needing more training time.](https://ss.rapidrecap.app/screens/pVs4XLk87oo/00-05-55.png)
![Screenshot at 08:24: A visual comparison highlighting the massive jump in performance \(29.7% lead over Llama 3.23D\) achieved by Instella 3B.](https://ss.rapidrecap.app/screens/pVs4XLk87oo/00-08-24.png)
