# The Open Source AI Model Beating GPT-5 on Agentic Performance

Source: https://www.youtube.com/watch?v=fblzzgnhZo4
Recap page: https://rapidrecap.app/video/fblzzgnhZo4
Generated: 2025-11-12T01:08:36.748+00:00

---
## Quick Overview

The open-source AI model Kimi K2 Thinking is outperforming proprietary models like GPT-5 and Claude Sonnet in agentic tasks, demonstrating a significant shift in AI development where Chinese open-source models are gaining ground on US counterparts by offering superior performance at a fraction of the cost, as evidenced by its performance on benchmarks and its ability to run locally.

**Key Points:**
- Kimi K2 Thinking scored 91% on the Humanity's Last Exam, surpassing GPT-5 and every other model mentioned, including DeepSeek V3, at a cost of $0.14/$0.29 per million tokens.
- The open-source lag is now measured in months, not years, with DeepSeek R1 following in 4 months, and GPT-5's lead over Kimi K2 shrinking to just 3 months.
- The 'closed model advantage window' has collapsed from 18+ months to 3-4 months, indicating a rapid catch-up by open-source models.
- The cost advantage is significant: Kimi K2 costs $0.16/$2.50 per million tokens compared to GPT-5's $1.25/$10.00 per million tokens for similar performance levels.
- The model can generate a full novel from one prompt, running up to 300 sequential tool calls per session, demonstrating advanced agentic capabilities.
- The release of Kimi K2 Thinking and subsequent models like MiniMax M2 suggests that the era of 'bigger models' is over, replaced by an era of 'smarter inference' and cost-effective local deployment.
- The video highlights that Chinese models are now competitive, with MiniMax's M2 ranking high on OpenRouter and Cognition AI's coding agent based on a Chinese model.

![Screenshot at 00:00: Announcement tweet from Kim.ai detailing the open-source Kimi K2 Thinking agent model's superior performance metrics \(SOTA on HLE and BrowseComp\) and low cost compared to competitors.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-00-00.png)

**Context:** The video discusses the rapidly evolving landscape of AI model development, focusing heavily on the competitive advancements made by Chinese open-source AI models challenging established US leaders like OpenAI and Anthropic. Key figures like Jensen Huang (Nvidia CEO) and various AI researchers/developers on X (formerly Twitter) are referenced to illustrate the closing performance and cost gap, particularly in agentic capabilities and coding benchmarks.

## Detailed Analysis

The discussion centers on the rapid progress of Chinese open-source AI models, specifically Kimi K2 Thinking from Moonshot AI, which is presented as outperforming proprietary models like GPT-5 and Claude Sonnet, especially in agentic tasks. Kimi K2 achieved a 91% score on Humanity's Last Exam, while DeepSeek R1, an earlier model, followed in four months, shrinking the 'closed model advantage window' to just 3-4 months. The cost difference is stark: Kimi K2 is significantly cheaper for comparable performance. Furthermore, the development of agents capable of running complex tasks (like generating a novel with up to 300 tool calls) locally on consumer hardware (M3 Ultras) is highlighted as a major advancement, signaling the end of the era dominated by massive, expensive models. A secondary point discusses how Chinese companies are betting on this cost-efficiency to win over global developers, even as US companies maintain advantages in hardware access. The overall conclusion is that the AI race is shifting, driven by cheap, efficient, and capable open-source models from China.

### Kimi K2 Thinking Release & Performance

- SOTA on HLE (94.4%) and BrowseComp (60.2%)
- Executes up to 200-300 sequential tool calls without human interference
- Excels in reasoning, agentic search, and coding.

### Cost & Efficiency Benchmarks

- Kimi K2 costs $0.14/$0.29 per million tokens (DeepSeek V3 comparison), significantly cheaper than GPT-5 ($1.25/$10.00 per million tokens).

### The Open-Source Lag

- Closed model advantage window collapsed from 18+ months to 3-4 months; China is treating AI like EV manufacturing, focusing on price/accessibility.

### Agentic Capabilities

- Agents can run for 1h:30m, demonstrating complex, multi-step reasoning in recent memory.

### Future Predictions (Bindu Reddy)

- 2026 will be the year of open weights; US labs will face competition; DeepSeek R2 and SOTA image/video models are coming.

### Developer Mindset Shift

- Entering the 'good enough and cheap' era; 80% of AI use cases solved by $2.50/million-token models; AI startup defensibility is based on using the best model, not necessarily building one.

![Screenshot at 00:00: Announcement tweet from Kim.ai detailing the open-source Kimi K2 Thinking agent model's superior performance metrics \(SOTA on HLE and BrowseComp\) and low cost compared to competitors.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-00-00.png)
![Screenshot at 00:03: Tweet from a Chinese source detailing US vs. China data center counts, suggesting China is building rapidly, fueling geopolitical competition.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-00-03.png)
![Screenshot at 00:20: Jensen Huang quote emphasizing China's nanosecond lead in AI and the need for the US to race ahead.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-00-20.png)
![Screenshot at 00:39: Rohan Paul tweet highlighting Kimi K2's superior performance over GPT-5, Cloud Sonnet 4.5, and Grok 4 on agentic tool use.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-00-39.png)
![Screenshot at 00:50: Rohan Paul tweet detailing Kimi K2's performance at commodity prices and the shift away from massive models to 'smarter inference'.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-00-50.png)
![Screenshot at 01:31: Tweet from Deedy contrasting the open-source lag \(months, not years\) and the collapse of the 'closed model advantage window' to 3-4 months.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-01-31.png)
![Screenshot at 03:14: Kim.ai tweet showing K2 Thinking is now live on Kimi.ai in chat mode, accessible via API, featuring bar graphs comparing performance.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-03-14.png)
![Screenshot at 03:44: Deedy tweet showing Kimi K2 scoring 91% on Humanity's Last Exam, outperforming GPT-5 and others, and running 15 tokens/sec on an M3 Ultra.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-03-44.png)
![Screenshot at 04:14: Side-by-side comparison showing Chinese models are now competitive with US models on coding benchmarks.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-04-14.png)
![Screenshot at 05:14: Machina tweet noting K2 beating Gemini 3 and suggesting Google's advantage in data is being eroded by capable, smaller teams building open-source models.](https://ss.rapidrecap.app/screens/fblzzgnhZo4/00-05-14.png)
