# Ben Hates Theo's Tierlist

Source: https://www.youtube.com/watch?v=qGH7X1gdv6k
Recap page: https://rapidrecap.app/video/qGH7X1gdv6k
Generated: 2026-08-29T03:17:43.474+00:00

---
## The Gist

Ben and Theo battle through a rigorous, high-stakes AI model tier list, ultimately placing Fable 5 and GPT-5.6-Sol at the absolute peak in S+ while demoting Gemini 3.1 Pro and Muse Spark 1.2. The hosts ruthlessly evaluate 16 frontier models based on reasoning, speed, pricing, and practical coding workflows.

## Quick Overview

Ben and Theo debate and finalize an exhaustive AI model tier list across four hours of intense discussion, grading 16 frontier models from S+ down to F. Fable 5 and GPT-5.6-Sol claim the S+ tier due to elite reasoning and deep reasoning efficiency, while Gemini 3.1 Pro and Muse Spark 1.2 anchor the lower rungs. The hosts meticulously factor in API pricing drops, raw token generation speeds, multimodal capabilities, and real-world coding agent benchmarks.

**Key Points:**
- Fable 5 and GPT-5.6-Sol secure the S+ tier because their exceptional reasoning and coding capabilities outweigh their steep pricing.
- Gemini 3.1 Pro is placed into the Google tier at the very bottom of the list due to severe limitations in real-world agentic workflows.
- Grok 4.6 lands firmly in the B tier, driven by its solid coding performance and strong subscription value inside Cursor.
- GLM-5.3-Flash earns a spot in the D tier, praised for its incredible frontier intelligence at a shockingly low price point of 4.5 cents per task.
- OpenAI recently slashed API pricing for GPT-5.6-Sol by twenty percent over a three-month period to aggressively compete with Anthropic.
- Muse Spark 1.2 is placed in the B tier, offering incredible value and blistering output speeds for specific coding and metadata tasks.
- Opus 5 drops down to the F tier because its high hallucination rate, lack of practical coding utility, and expensive pricing make it unusable for developers.

![Screenshot at 00:00: The final AI model tier list reveals Fable 5 and GPT-5.6-Sol at the top in S+, while Gemini 3.1 Pro sits at the very bottom.](https://ss.rapidrecap.app/screens/qGH7X1gdv6k/00-00-00.jpg)

**Context:** Ben and Theo host Nerdsnipe, a developer-focused podcast analyzing frontier artificial intelligence models. In this episode, the two clash over a comprehensive model tier list, grading everything from OpenAI and Anthropic flagships to open weights and Chinese labs based on grueling real-world developer benchmarks.

## Detailed Analysis

Ben and Theo spend hours hammering out the definitive developer tier list for 16 major artificial intelligence models, breaking down performance metrics, agentic benchmarks, API costs, and personal frustrations. Throughout the episode, the hosts spar over edge cases, debating whether open-weight models like Grok 4.6 or budget powerhouses like GLM-5.3-Flash deserve higher placements than bloated enterprise offerings like Opus 5. They meticulously weigh the real-world utility of vision capabilities, prompt adherence, and long-horizon reinforcement learning frameworks. Ultimately, Fable 5 and GPT-5.6-Sol take the crown in S+ for unmatched coding prowess, while Gemini 3.1 Pro and Opus 5 face harsh demotions for failing to justify their high costs and inconsistent outputs.

### Fable 5 and GPT-5.6-Sol

The two undisputed champions of the tier list, capturing the S+ tier through sheer engineering dominance.

- Fable 5 operates like a wise owl, offering deep conversational intelligence and unmatched agentic code merging capabilities.
- GPT-5.6-Sol functions like a ruthless rottweiler that grips complex problems by the throat and refuses to let go until solved.
- Both models justify their high API costs by delivering pristine, mergeable code on complex, multi-step repositories.

![Screenshot at 148:33: Fable 5 and GPT-5.6-Sol are locked into the S+ tier after extensive debate between Ben and Theo.](https://ss.rapidrecap.app/screens/qGH7X1gdv6k/00-148-33.jpg)

### Luna and Ox Alpha

The upper-tier powerhouses that balance extreme cost efficiency with top-tier capability.

- Luna receives an A tier placement for offering absurdly cheap API pricing that undercuts almost all Western competition.
- Ox Alpha drops with a massive one million context window, multi-modal features, and zero data retention guarantees.
- Both models provide elite developer utility without breaking the bank on high-volume agentic tasks.

![Screenshot at 149:14: Luna and Ox Alpha occupy the A tier due to their stellar pricing and massive context capabilities.](https://ss.rapidrecap.app/screens/qGH7X1gdv6k/00-149-14.jpg)

### DeepSeek v4 Flash and Muse Spark 1.2

The high-performing middle tier that bridges the gap between speed and budget constraints.

- DeepSeek v4 Flash lands in the B tier due to blistering token generation speeds and strong localized deployment options.
- Muse Spark 1.2 punches well above its weight class, providing incredible metadata and title generation performance for pennies.
- Both models excel in environments where latency is prioritized over exhaustive multi-step reasoning.

![Screenshot at 150:18: DeepSeek v4 Flash and Muse Spark 1.2 are slotted into the B tier for exceptional speed and value.](https://ss.rapidrecap.app/screens/qGH7X1gdv6k/00-150-18.jpg)

### Grok 4.6 and Kimi k3

Reliable workhorses that serve as comfortable everyday tools for practical developers.

- Grok 4.6 earns a C tier spot because its integration inside Cursor and solid reasoning make it a dependable daily driver.
- Kimi k3 features impressive 3D spatial capabilities and strong writing chops, though it suffers from slow generation times.
- Both models represent the middle of the pack, offering great utility for developers willing to work around minor quirks.

![Screenshot at 151:12: Grok 4.6 and Kimi k3 settle into the C tier as dependable mid-tier developer alternatives.](https://ss.rapidrecap.app/screens/qGH7X1gdv6k/00-151-12.jpg)

### GLM-5.3-Flash and DeepSeek v4 Pro

The budget contenders that deliver surprising intelligence despite minor architectural limitations.

- GLM-5.3-Flash delivers frontier-level performance at a meager cost of 4.5 cents per task, impressing both hosts.
- DeepSeek v4 Pro handles heavy coding tasks well, but its lack of native vision support holds it back from higher tiers.
- Both models prove that open-weight and alternative labs are rapidly closing the gap with Western giants.

![Screenshot at 152:36: GLM-5.3-Flash and DeepSeek v4 Pro are placed in the D tier for offering budget performance with minor drawbacks.](https://ss.rapidrecap.app/screens/qGH7X1gdv6k/00-152-36.jpg)

### Opus 5 and Sonnet 5

The disappointing heavyweights that failed to live up to their initial hype among developers.

- Opus 5 is relegated to the F tier because its high hallucination rate and over-complicated code edits frustrate developers.
- Sonnet 5 lands in the F tier due to skyrocketing API costs that outpace its actual reasoning improvements.
- Both models suffer from severe regressions in developer sentiment compared to their previous iterations.

![Screenshot at 153:21: Opus 5 and Sonnet 5 hit rock bottom in the F tier due to high costs and frustrating hallucinations.](https://ss.rapidrecap.app/screens/qGH7X1gdv6k/00-153-21.jpg)

### Gemini 3.1 Flash and Gemini 3.1 Pro

Google's flagship offerings that miss the mark on practical coding agent benchmarks.

- Gemini 3.1 Flash achieves fast output speeds but falters on complex multi-file software engineering tasks.
- Gemini 3.1 Pro sits alone in the Google tier at the very bottom, criticized for poor agentic integration and weak context tracking.
- Both models require significant improvement before competing seriously with OpenAI and Anthropic flagships.

![Screenshot at 154:55: Gemini models occupy the lowest positions, with Gemini 3.1 Pro sitting at the bottom Google tier.](https://ss.rapidrecap.app/screens/qGH7X1gdv6k/00-154-55.jpg)

