# OpenAI Just Broke The Industry

Source: https://www.youtube.com/watch?v=NyW7EDFmWl4
Recap page: https://rapidrecap.app/video/NyW7EDFmWl4
Generated: 2025-08-05T20:01:32.69+00:00

---
## Quick Overview

OpenAI has released two new open-weight language models, gpt-oss-120b and gpt-oss-20b, which offer competitive performance on reasoning and coding tasks, are optimized for efficient deployment, and are available under the permissive Apache 2.0 license.

**Key Points:**
- OpenAI released two new open-weight language models, gpt-oss-120b and gpt-oss-20b, under the Apache 2.0 license.
- These models demonstrate strong performance on reasoning, coding, and tool-use tasks, competing with or exceeding similarly sized proprietary models.
- gpt-oss-120b runs efficiently on an 80GB GPU, while gpt-oss-20b requires only 16GB, making it suitable for on-device applications.
- The models are trained using a mix of reinforcement learning and advanced techniques, focusing on reasoning, efficiency, and real-world usability.
- They show competitive results across various benchmarks including Codeforces, Humanity's Last Exam, MMLU, and HealthBench.
- OpenAI emphasizes safety and has implemented rigorous training and evaluation processes.
- The release aims to lower barriers for AI adoption, foster innovation, and expand access to powerful AI tools globally.

![Screenshot at 00:10: Bar chart showing the Elo ratings of various language models on Codeforces competition code, highlighting the strong performance of gpt-oss-120b with tools at 2622, close to the top-performing o4-mini with tools at 2706.](https://ss.rapidrecap.app/screens/NyW7EDFmWl4/00-00-10.png)

**Context:** OpenAI has recently released two new open-weight language models, gpt-oss-120b and gpt-oss-20b. This release is significant as it aligns with the US AI Action Plan's focus on strengthening American open-source AI foundations and represents a broader trend towards making advanced AI more accessible. The models are designed to be efficient, performant, and available under a permissive license, aiming to empower a wider range of users and foster innovation in the AI community.

## Detailed Analysis

OpenAI announces the release of two new open-weight language models, gpt-oss-120b and gpt-oss-20b, aiming to push the frontier of open-weight reasoning models. These models deliver strong real-world performance at a low cost, are available under the Apache 2.0 license, and outperform similarly sized open models on reasoning tasks and tool use capabilities. They are optimized for efficient deployment on consumer hardware and were trained using a mix of reinforcement learning and techniques informed by OpenAI's internal models like o3 and other frontier systems. The gpt-oss-120b model achieves near-parity with OpenAI's o4-mini on core reasoning benchmarks while running efficiently on an 80GB GPU. The gpt-oss-20b model delivers similar results to OpenAI's o3-mini on common benchmarks and can run on edge devices with just 16GB of memory, making it ideal for on-device use cases, local inference, and rapid iteration without costly infrastructure. Both models also perform strongly on tool use, few-shot function calling, Chain-of-Thought (CoT) reasoning, and HealthBench, with some performance even exceeding proprietary models like o1 and GPT-4o. They are compatible with OpenAI's Responses API and designed for agentic workflows, including web search, Python code execution, and reasoning capabilities. OpenAI emphasizes safety as foundational, conducting thorough safety training and evaluations, including testing adversarially fine-tuned versions of the models. The models are trained on a mostly English, text-only dataset focusing on STEM, coding, and general knowledge, using a tokenizer that is also open-sourced. OpenAI is excited to provide these models to empower developers, enterprises, and governments, fostering innovation and broader access to AI capabilities.

### Introduction

- Release of gpt-oss-120b and gpt-oss-20b
- Open-weight models
- Apache 2.0 license
- High performance on reasoning and tool use
- Efficient deployment
- Training methodology
- Benchmarking results

### Model Performance

- gpt-oss-120b vs. o4-mini on reasoning benchmarks
- gpt-oss-20b vs. o3-mini on common benchmarks
- On-device capabilities with 16GB memory
- Tool use, few-shot function calling, CoT reasoning, HealthBench performance

### Training and Architecture

- Transformer architecture with mixture-of-experts
- Parameter counts: 120b (5.1B active), 20b (3.6B active)
- Attention patterns
- Rotary Positional Embedding (RoPE)
- Context lengths up to 128k

### Dataset and Tokenizer

- Mostly English, text-only dataset
- Focus on STEM, coding, general knowledge
- Open-sourced tokenizer (o200k_harmony)

### Post-training

- Similar process to o4-mini
- Supervised fine-tuning stage
- High-compute RL stage
- Alignment with OpenAI Model Spec
- CoT reasoning and tool use training

### Reasoning Efforts

- Support for low, medium, and high reasoning efforts
- Trade-off between latency and performance
- System message for setting reasoning effort

### Evaluations

- Standard academic benchmarks
- Comparison with other OpenAI models (o3, o3-mini, o4-mini)
- Performance on Codeforces, Humanity's Last Exam, HealthBench, AIME, Tau-Bench

### Availability and Partnerships

- Freely available on Hugging Face
- Natively quantized in MXFP4
- Harmony prompt format and renderer
- Reference implementations for PyTorch, Apple Metal
- Partnerships with Azure, Hugging Face, vLLM, Ollama, etc.
- GPU-optimized version for Windows devices
- Open-source AI foundations and US Action Plan alignment
- Open-source ecosystem benefits

![Screenshot at 00:10: Bar chart comparing Elo ratings for various language models on Codeforces competition code, showing gpt-oss-120b \(with tools\) at 2622 and o4-mini \(with tools\) at 2706.](https://ss.rapidrecap.app/screens/NyW7EDFmWl4/00-00-10.png)
![Screenshot at 01:06: Text excerpt introducing gpt-oss-120b and gpt-oss-20b as state-of-the-art open-weight language models.](https://ss.rapidrecap.app/screens/NyW7EDFmWl4/00-01-06.png)
![Screenshot at 07:46: Bar chart showing accuracy \(%\) on 'Humanity's Last Exam' for different models, with gpt-oss-120b \(with tools\) at 19% and o3 \(with tools\) at 24.9%.](https://ss.rapidrecap.app/screens/NyW7EDFmWl4/00-07-46.png)
![Screenshot at 08:31: Bar charts comparing accuracy \(%\) on AIME 2024 and AIME 2025 \(competition math\) for various models.](https://ss.rapidrecap.app/screens/NyW7EDFmWl4/00-08-31.png)
![Screenshot at 08:58: Bar charts comparing accuracy \(%\) on GPQA Diamond \(without tools\) and MMLU \(questions across academic disciplines\) for different models.](https://ss.rapidrecap.app/screens/NyW7EDFmWl4/00-08-58.png)
![Screenshot at 09:13: Bar chart showing accuracy \(%\) on Tau-Bench Retail \(function calling\) for gpt-oss-120b \(67.8%\), o3 \(70.4%\), and o4-mini \(65.6%\).](https://ss.rapidrecap.app/screens/NyW7EDFmWl4/00-09-13.png)
![Screenshot at 10:07: Text section discussing monitoring reasoning model's Chain-of-Thought \(CoT\) for detecting misbehavior and the importance of non-supervised CoT.](https://ss.rapidrecap.app/screens/NyW7EDFmWl4/00-10-07.png)
![Screenshot at 12:18: Tweet from Clement Delangue expressing excitement about the OSS-GPT release and its implications for open-source AI.](https://ss.rapidrecap.app/screens/NyW7EDFmWl4/00-12-18.png)
![Screenshot at 13:07: Graphic showing 'OpenAI' and 'HF' logos, symbolizing the partnership.](https://ss.rapidrecap.app/screens/NyW7EDFmWl4/00-13-07.png)
