# Alibaba is going all in on Qwen…

Source: https://www.youtube.com/watch?v=SquU4Bpc73Y
Recap page: https://rapidrecap.app/video/SquU4Bpc73Y
Generated: 2025-10-03T16:32:34.993+00:00

---
## Quick Overview

Alibaba Cloud has aggressively entered the AI race by announcing its Qwen3-Max model at the Apsara Conference, positioning itself for full-stack AI dominance with a model boasting over one trillion parameters trained on 36 trillion tokens, while also open-sourcing Qwen3-VL, which achieves state-of-the-art results in multimodal reasoning benchmarks, outperforming competitors like GPT-5 Chat and Gemini 2.5 Pro on certain tasks.

**Key Points:**
- Alibaba Cloud announced the Qwen3-Max model, featuring over one trillion parameters and pre-trained on 36 trillion tokens, aiming for full-stack AI dominance (0:04, 1:25).
- The Qwen3-Max-Instruct variant achieved a 39.4% accuracy on a benchmark, ranking #1 above GPT-5 Chat (32.8%) and Gemini 2.5 Pro (18.9%) (2:06).
- The Qwen team also released Qwen3-VL, a vision-language model where the instruct version matches or exceeds Gemini 2.5 Pro on major visual perception benchmarks (1:45).
- Qwen3-Omni, a multimodal model, demonstrated capabilities across four dimensions: Smarter (reasoning), Multilingual, Faster (low latency), and Longer (context length) (2:14).
- The video highlights a three-phase roadmap to ASI: Stage 1 (Emergent Intelligence/Learning from Humans), Stage 2 (AI Agency/Autonomous action), and Stage 3 (Self-improvement/Self-iteration) (0:41).
- Alibaba is open-sourcing flagship models like Qwen3-VL-235B-A22B, emphasizing an 'Open-Source & Open-Weight' strategy (2:12).
- The video encourages viewers to learn coding principles through interactive lessons on Brilliant, focusing on problem-solving over memorization (2:33, 2:42).

![Screenshot at 1:25: The Qwen3-Max-Base model size is revealed, noting it has over 1 trillion parameters and was trained on 36 trillion tokens, emphasizing its massive scale.](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-01-25.png)

**Context:** The video provides a news update and analysis of recent major announcements from Alibaba Cloud, specifically focusing on their Qwen large language model family, unveiled during their Apsara Conference. The context is set against the backdrop of intense global competition in AI, particularly involving tech giants like OpenAI and Google, as Alibaba aggressively pushes its models for performance, size, and open-sourcing capabilities to challenge established leaders.

## Detailed Analysis

Alibaba Cloud declared its intent for full-stack AI dominance by unveiling the Qwen3-Max model during the Apsara Conference (0:04). This flagship model, Qwen3-Max-Base, features over one trillion parameters and was pre-trained on a massive 36 trillion tokens (1:25). The instructional version, Qwen3-Max-Instruct, demonstrated superior performance on several benchmarks, achieving 39.4% accuracy on one test, surpassing GPT-5 Chat (32.8%) and Gemini 2.5 Pro (18.9%) (2:06). The company also introduced Qwen3-VL, a vision-language model designed for sharper vision and deeper thought, which achieves state-of-the-art results in multimodal reasoning and visual perception tasks, with its Instruct version matching or exceeding Gemini 2.5 Pro (1:45). The presentation detailed a 3-phase roadmap to Artificial Super Intelligence (ASI): Stage 1 involves learning from humans (Emergent Intelligence), Stage 2 focuses on AI Agency (Autonomous Action), and Stage 3 is Self-Iteration using real-world data for self-learning (0:41). Furthermore, Alibaba is committed to an 'Open-Source & Open-Weight' approach, releasing models like Qwen3-VL-235B-A22B (2:12). The video concludes by promoting Brilliant.org for learning the necessary math and computer science foundations required to understand these advanced AI concepts, emphasizing interactive problem-solving over passive learning (2:33).

### Alibaba's AI Ambition

- Alibaba Cloud targets full-stack AI dominance with Qwen3-Max launch
- Qwen3-Max-Base exceeds 1 trillion parameters and 36 trillion tokens training data
- A 3-phase roadmap to ASI (Learning, Agency, Self-iteration) is outlined (0:04
- 1:25
- 0:41)

### Qwen3-Max Performance

- Qwen3-Max-Instruct ranks #1 on a key leaderboard with 39% accuracy, beating GPT-5 Chat and Gemini 2.5 Pro
- Qwen3-Max-Thinking (Heavy) shows competitive reasoning scores against Grok-4 and GPT-5 Pro (2:06
- 1:40)

### Qwen3-VL Release

- New vision-language model released, open-sourced in both Instruct and Thinking versions
- Qwen3-VL-235B-A22B Instruct version matches or exceeds Gemini 2.5 Pro on visual perception benchmarks (1:45)

### Qwen3-Omni Capabilities

- Multimodal model excelling in Smarter (reasoning), Multilingual tasks, Faster processing (234ms latency), and Longer context handling (2:14)

### Openness Strategy

- Alibaba emphasizes open-sourcing flagship models like Qwen3-VL, promoting an 'Open-Source & Open-Weight' ethos (2:12)

### Sponsor Segment (Brilliant)

- Brilliant offers courses on Solving Equations, Thinking in Code, Scientific Thinking, and Programming with Python
- Focus is on interactive, problem-solving lessons rather than lectures (2:33
- 2:42)

![Screenshot at 0:02: Meme showing tech CEOs \(Zuckerberg, Musk, Ellison, unknown\) engulfed in flames, representing the competitive pressure in the AI space.](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-00-02.png)
![Screenshot at 0:05: A presenter at the APSARA Conference detailing the Qwen knowledge engineering workflow for financial datasets.](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-00-05.png)
![Screenshot at 0:10: Propaganda-style art depicting a Chinese astronaut triumphing over US astronauts, symbolizing China's perceived lead in the space/tech race.](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-00-10.png)
![Screenshot at 0:25: A tweet meme humorously illustrating the rapid succession of Qwen model releases \(Qwen1, 2.5, 3, etc.\).](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-00-25.png)
![Screenshot at 0:41: Alibaba's 3-Phase Roadmap to ASI: Stage 1 \(Emergent Intelligence\), Stage 2 \(AI Agency\), Stage 3 \(Self-improvement/Outperforming Humans\).](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-00-41.png)
![Screenshot at 0:55: Bar chart comparing coding accuracy \('With thinking' vs 'Without thinking'\) for Wide Coders, Staff Engineers, and Engineering Graduates, used to frame AI agency capabilities.](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-00-55.png)
![Screenshot at 1:36: Text Arena leaderboard showing Qwen3-Max-Preview ranking 3rd, ahead of several leading proprietary models.](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-01-36.png)
![Screenshot at 1:40: Bar chart comparing Qwen3-Max-Instruct performance across multiple benchmarks \(SuperGPQA, AIME25, LiveCodeBench v6, etc.\) against other models like Claude Opus 4 and DeepSeek V3.1.](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-01-40.png)
![Screenshot at 2:14: Infographic detailing the four key aspects of the Qwen3-Omni multimodal model: Smarter, Multilingual, Faster, and Longer context.](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-02-14.png)
![Screenshot at 2:57: Brilliant.org call to action slide with a QR code and URL, promoting learning skills like coding and scientific thinking.](https://ss.rapidrecap.app/screens/SquU4Bpc73Y/00-02-57.png)
