# Qwen 3.5 - The next NEXT model

Source: https://www.youtube.com/watch?v=X6mL3cdPiCg
Recap page: https://rapidrecap.app/video/X6mL3cdPiCg
Generated: 2026-02-17T12:03:30.445+00:00

---
## Quick Overview

The Qwen 3.5 release introduces the Qwen3.5-397B-A17B flagship vision-language model, which is open-sourced, utilizes an innovative hybrid architecture combining Linear Attention and Sparse Mixture-of-Experts (MoE), and demonstrates significant performance leaps in Coding and Multimodal Understanding benchmarks compared to previous Qwen models and competitors like Claude Opus 4.5 and Gemini 3 Pro, while achieving 8.6x to 19.0x inference throughput boost over Qwen-Max due to only activating 17 billion out of 397 billion total parameters per forward pass.

**Key Points:**
- The flagship model Qwen3.5-397B-A17B is now open-sourced, featuring 397 billion total parameters with only 17 billion activated per forward pass.
- The model achieves a significant performance leap in Coding and Multimodal Understanding benchmarks, often surpassing competitors like Claude Opus 4.5 and Gemini 3 Pro on several metrics (00:07, 00:08, 01:28, 01:37).
- Inference efficiency sees a major boost: Qwen3.5-397B-A17B achieves 8.6x decode throughput at 32K context length and 19.0x at 256K context length compared to Qwen-Max (00:17).
- Qwen 3.5 employs an innovative hybrid architecture fusing Linear Attention (via Gated Delta Networks) with a Sparse Mixture-of-Experts (MoE) structure (00:12).
- Multimodal capabilities are enhanced, supporting 201 languages/dialects, a 250K vocabulary, and showing 10-60% coding/decoding efficiency improvement across most languages (00:10, 06:00).
- The Qwen3-Next architecture is built upon higher-sparsity MoE, Gated DeltaNet, and Gated Attention, leading to performance comparable to the 1T-parameter Qwen3-Max Base model despite having fewer active parameters (00:07, 04:53).
- The model is accessible via Qwen Chat (chat.qwen.ai) with modes for Auto, Thinking, and Fast processing, and advanced features like reasoning and web search are enabled via parameters in the ModelStudio (08:11, 08:37).

![Screenshot at 00:41: The Qwen 3.5 Open-source Release graphic highlights the four key pillars: Inference Efficiency, Hybrid Architecture, Native Multimodality, and Global Scale, symbolized by the Qwen bear mascot holding a key to a treasure chest glowing with multimodal data visualizations.](https://ss.rapidrecap.app/screens/X6mL3cdPiCg/00-00-41.jpg)

**Context:** This video discusses the open-source release of the Qwen 3.5 series of large language models, focusing primarily on the flagship Qwen3.5-397B-A17B model, which is positioned as a native vision-language agent. The presentation highlights architectural improvements, significant performance gains across various benchmarks (especially in coding and multimodal tasks), and massive gains in inference efficiency achieved through a novel MoE architecture.

## Detailed Analysis

The video announces the open-source release of Qwen 3.5, featuring the flagship Qwen3.5-397B-A17B model. This model is native vision-language and demonstrates outstanding results across reasoning, coding, agent capabilities, and multimodal understanding. Key to its performance is the innovative hybrid architecture combining Linear Attention (via Gated Delta Networks) with a Sparse Mixture-of-Experts (MoE). Although the model has 397 billion total parameters, only 17 billion are activated per forward pass, leading to remarkable inference efficiency. Benchmarks show Qwen3.5-397B-A17B matching or beating the performance of the >1T-parameter Qwen3-Max Base model. Inference throughput is significantly boosted, showing an 8.6x improvement at 32K context length and a 19.0x improvement at 256K context length compared to Qwen-Max. Furthermore, the model expands its multilingual support from 119 to 201 languages/dialects, which is noted as a crucial decision for competitive modeling. The presentation also shows Qwen3.5-397B-A17B outperforming competitors like GPT-5.2 High, Claude Opus 4.5, and Gemini 3 Pro in environment scaling tests (07:07). Finally, the video details access via chat.qwen.ai, offering Auto, Thinking, and Fast modes, and highlights the ModelStudio where advanced features like reasoning and web search can be enabled using parameters like 'enable_thinking' and 'enable_search'. Demos showcase its ability to handle complex tasks like spreadsheet manipulation and generating slide decks from image prompts.

### Qwen 3.5 Release & Architecture

- Flagship model Qwen3.5-397B-A17B open-sourced
- Hybrid architecture uses Linear Attention (Gated DeltaNet) + Sparse MoE
- 397B total parameters, only 17B activated per forward pass (00:04, 00:12, 04:45)

### Performance Gains

- Significant leap in Coding and Multimodal Understanding benchmarks
- Qwen3.5-397B-A17B matches or beats >1T-parameter Qwen3-Max Base (00:07, 01:28)

### Inference Efficiency

- 8.6x throughput boost at 32K context length; 19.0x boost at 256K context length compared to Qwen-Max (00:17)

### Multimodality & Language

- Natively multimodal via early text-vision fusion
- Multilingual coverage grows from 119 to 201 languages/dialects
- 250K vocabulary (05:58, 06:00)

### Scaling & Comparison

- Qwen3.5-397B-A17B Thinking outperforms Claude Opus 4.5, Gemini 3 Pro, and Qwen3-Max Thinking in environment scaling tests (07:07)

### Agent Capabilities & Access

- Model supports Think, Search, Use Tools, and Build in multimodal context via Qwen Chat (chat.qwen.ai)
- ModelStudio enables advanced reasoning via 'enable_thinking' and web search via 'enable_search' parameters (08:11, 08:37)

### Demo Examples

- Showcase of complex tasks including spreadsheet manipulation (SUM formula entry) and generating a multi-slide presentation from an image prompt (01:20, 08:22, 10:22)

![Screenshot at 00:08: Qwen 3.5 benchmark results showing Qwen3.5-397B-A17B leading competitors across multiple multimodal tasks \(e.g., GPQA Diamond at 91.9\) \(00:08\)](https://ss.rapidrecap.app/screens/X6mL3cdPiCg/00-00-08.jpg)
![Screenshot at 00:17: Decode Throughput comparison showing Qwen3.5-397B-A17B achieving 8.6x throughput boost at 32K context length over Qwen-Max \(00:17\)](https://ss.rapidrecap.app/screens/X6mL3cdPiCg/00-00-17.jpg)
![Screenshot at 00:41: The main launch graphic summarizing Qwen 3.5's key features: Inference Efficiency, Hybrid Architecture, Native Multimodality, and Global Scale \(00:41\)](https://ss.rapidrecap.app/screens/X6mL3cdPiCg/00-00-41.jpg)
![Screenshot at 01:11: Hugging Face repository page showing the Qwen3-Next models, including the 80B parameter version \(01:11\)](https://ss.rapidrecap.app/screens/X6mL3cdPiCg/00-01-11.jpg)
![Screenshot at 07:08: Average Ranking vs. Environment Scaling chart illustrating Qwen3.5-397B-A17B Thinking achieving a top rank against GPT-5.2 High, Claude Opus 4.5, and Gemini 3 Pro after extensive training environments \(07:07\)](https://ss.rapidrecap.app/screens/X6mL3cdPiCg/00-07-08.jpg)
