# My Review on Kimi K2 Thinking After Days of Testing…

Source: https://www.youtube.com/watch?v=fjCwC6BYKAE
Recap page: https://rapidrecap.app/video/fjCwC6BYKAE
Generated: 2025-11-13T15:33:32.203+00:00

---
## Quick Overview

Moonshot AI's Kimi K2 Thinking model demonstrates superior performance in agentic tasks, outperforming top models like GPT-5 and Claude 4.5 on the r²-Bench Telecom test with a 93% score, though it was slower and had minor implementation issues (like UI theme mismatch in one test) compared to Claude, ultimately proving its strong potential in complex reasoning and coding capabilities.

**Key Points:**
- Kimi K2 Thinking achieved the top score of 93% on the r²-Bench Telecom (Agentic Tool Use) benchmark, surpassing GPT-5 (87%) and Claude 4.5 Sonnet (78%).
- In a complex coding test to build a 3D pinball game, Kimi K2 took longer than Claude but successfully implemented the game, including sound effects, which Claude failed to do.
- In a project management authentication test, Kimi K2 successfully implemented Firebase authentication, while Claude generated a login/signup page that mismatched the website's dark UI theme.
- Kimi K2 Thinking is priced significantly lower than competitors, with an output price of $0.60 per 1M tokens for the k2-thinking model compared to Claude Sonnet 4.5's $3 per 1M input tokens.
- Kimi K2 Thinking can execute up to 200-300 sequential tool calls without human interference, highlighting its advanced reasoning capabilities.
- The video also showcases the capabilities of Make.com for visual orchestration of AI agents and automations across various business functions.

![Screenshot at 00:17: Kimi K2 Thinking achieving 93% on the r²-Bench Telecom \(Agentic Tool Use\) chart, placing it at the top of the benchmark against competitors like GPT-5 and Claude models.](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-00-17.png)

**Context:** This video reviews and benchmarks the Kimi K2 Thinking model from Moonshot AI, a Chinese AI company, comparing its performance, speed, and cost-efficiency against leading models like OpenAI's GPT-5 and Anthropic's Claude 4.5 (specifically Sonnet and Opus) across several agentic and coding tasks. The review aims to strip away marketing hype to assess real-world capabilities.

## Detailed Analysis

The video performs three main tests to evaluate Kimi K2 Thinking: agentic tool use, complex coding, and feature implementation (authentication). In agentic tool use (r²-Bench Telecom), Kimi K2 scored 93%, clearly leading GPT-5 (87%) and Claude models. For the first coding test (creating a 3D fashion website prototype), Kimi K2 completed the task but had minor UI bugs, while Claude took longer and failed to implement sound effects. For the second coding test (creating a 3D pinball game), Kimi K2 took longer than Claude but successfully implemented the game with sound effects, whereas Claude's implementation was not fully dynamic. For the authentication feature test, Kimi K2 successfully integrated Firebase authentication without breaking existing code, while Claude generated the required UI but failed to adhere to the existing dark theme. Cost-wise, Kimi K2's pricing is shown to be highly competitive, significantly cheaper than Claude's offerings. The video concludes that while Kimi K2 is incredibly impressive and cost-effective, it still has minor inconsistencies and speed issues compared to Claude, though it excels in agentic benchmarks. The latter part of the video pivots to showcase Make.com's visual orchestration platform for building and managing AI agent workflows.

### Kimi K2 Introduction

- Introduced as Moonshot AI's best open-source thinking model
- Built as a thinking agent capable of 200-300 sequential tool calls
- Claims major gains in reasoning, coding, writing, and general capabilities.

### Pricing Comparison

- Kimi k2-thinking output price is $0.60 per 1M tokens
- Claude Sonnet 4.5 input price is $3 per 1M tokens
- Kimi K2 is significantly cheaper than comparable models.

### Test 1

- Agentic Tool Use (r²-Bench Telecom): Kimi K2 Thinking scored 93%
- GPT-5 Codex (high) scored 87%
- Claude 4.5 Sonnet scored 78%
- Kimi K2 is confirmed as the top performer in this area.

### Test 2

- 3D Fashion Website Prototype: Kimi K2 successfully created the site with interactive 3D models and color changes
- Claude took longer and produced a site with broken UI elements like incorrect colors.

### Test 3

- 3D Pinball Game Recreation: Kimi K2 successfully recreated the game with sound effects
- Claude's version was static and lacked sound effects, failing to fully meet instructions.

### Test 4

- Project Management Auth Feature: Kimi K2 successfully added Firebase authentication
- Claude generated the login/signup UI but mismatched the dark theme of the existing site.

### Sponsor Segment (Make.com)

- Promotes Make.com for real-time visual orchestration of AI agents and automations
- Highlights ability to connect 3000+ apps and share workflows.

![Screenshot at 00:02: Introduction screen for Moonshot AI, showing options for Kimi and Kimi Open Platform.](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-00-02.png)
![Screenshot at 00:09: Moonshot AI pricing table highlighting the kimo-k2-thinking model's low cost \($0.60/1M output\).](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-00-09.png)
![Screenshot at 00:17: Bar chart showing Kimi K2 Thinking leading the r²-Bench Telecom \(Agentic Tool Use\) at 93%.](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-00-17.png)
![Screenshot at 00:24: Side-by-side cost comparison between Claude Sonnet 4 and Kimi K2 0905 Preview, illustrating Kimi's lower input and output pricing.](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-00-24.png)
![Screenshot at 00:37: VS Code interface showing the AI agent actively working on coding tasks for the 3D pinball recreation.](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-00-37.png)
![Screenshot at 00:49: Kimi web interface showing the prompt input area ready for testing the model.](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-00-49.png)
![Screenshot at 01:04: Test One screen indicating the start of the UI design capability test for the fashion website.](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-01-04.png)
![Screenshot at 01:17: Terminal output showing Claude running the Next.js setup command, ready for code generation.](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-01-17.png)
![Screenshot at 01:43: The 3D Fashion Website prototype generated by Kimi, featuring interactive 3D models.](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-01-43.png)
![Screenshot at 02:24: Context usage report showing Claude consuming 137k/200k tokens for the fashion website task, which is 68% of its context window, costing $0.2271 for the run.](https://ss.rapidrecap.app/screens/fjCwC6BYKAE/00-02-24.png)
