Instella: Fully Open Language Models with Stellar Performance

Quick Overview

Instella 3B achieves superior performance across benchmarks compared to similarly sized open and fully open models, notably outperforming Llama 3.5 by a significant margin on the GSM8K math reasoning benchmark (4.1% lead) and achieving the highest reported performance among all models in its size class on the GPQA benchmark, demonstrating that smaller, highly-tuned models can surpass larger, less transparent ones.

Key Points: Instella 3B outperforms Llama 3.5 by 4.1% on the GSM8K math reasoning benchmark. Instella 3B achieved the highest reported performance among all models in its size class on the GPQA benchmark. The model was trained using a two-stage process: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning from Human Feedback (RLHF) using Direct Preference Optimization (DPO). The training involved 128K instruction-response pairs and 58 Billion tokens of synthetic data generated using a larger teacher model. The training methodology focused on maximizing performance on complex reasoning tasks like Truthful QA (TQA) and General Knowledge Question Answering (GKQA). Instella 3B demonstrated superior performance compared to competitors like Llama 3.5, beating it by 14% on average across benchmarks. The model's 128K context window version achieved a 4.1% lead over Llama 3.5 on GSM8K.

Context: The video introduces Instella 3B, a new large language model (LLM) developed by AMD, positioning it as a highly capable, fully open model designed to challenge existing proprietary and open-source leaders. The discussion centers on its performance metrics, particularly against models like Llama, and the innovative, multi-stage training methodology used to achieve this high level of reasoning and factual accuracy with a relatively small parameter count.

Detailed Analysis

The Instella 3B model demonstrates stellar performance, especially for its size (3 billion parameters), setting a new standard for fully open language models. The key takeaway is that specialized training, emphasizing complex reasoning and safety alignment, allows smaller models to significantly outperform larger, less transparent competitors. AMD achieved this through a rigorous two-stage training process: Supervised Fine-Tuning (SFT) on 128K instruction-response pairs, followed by RLHF using Direct Preference Optimization (DPO) guided by human preferences. This process focused on teaching the model to reason logically, rather than just recall facts, and to refuse unsafe or nonsensical queries. Performance comparisons show Instella 3B beating Llama 3.5 by a substantial margin (e.g., 4.1% lead on GSM8K) and setting new state-of-the-art scores for its class on benchmarks like GPQA. The researchers engineered the training to focus on high-quality, diverse synthetic data generation using a larger model to create context-rich training examples, ensuring the model generalizes well to real-world tasks while maintaining a small footprint.

Raw markdown version of this recap