Ben Hates Theo's Tierlist | Nerd Snipe
The Gist
Ben and Theo battle through a rigorous, high-stakes AI model tier list, ultimately placing Fable 5 and GPT-5.6-Sol at the absolute peak in S+ while demoting Gemini 3.1 Pro and Muse Spark 1.2. The hosts ruthlessly evaluate 16 frontier models based on reasoning, speed, pricing, and practical coding workflows.
Quick Overview
Ben and Theo debate and finalize an exhaustive AI model tier list across four hours of intense discussion, grading 16 frontier models from S+ down to F. Fable 5 and GPT-5.6-Sol claim the S+ tier due to elite reasoning and deep reasoning efficiency, while Gemini 3.1 Pro and Muse Spark 1.2 anchor the lower rungs. The hosts meticulously factor in API pricing drops, raw token generation speeds, multimodal capabilities, and real-world coding agent benchmarks.
Key Points: Fable 5 and GPT-5.6-Sol secure the S+ tier because their exceptional reasoning and coding capabilities outweigh their steep pricing. Gemini 3.1 Pro is placed into the Google tier at the very bottom of the list due to severe limitations in real-world agentic workflows. Grok 4.6 lands firmly in the B tier, driven by its solid coding performance and strong subscription value inside Cursor. GLM-5.3-Flash earns a spot in the D tier, praised for its incredible frontier intelligence at a shockingly low price point of 4.5 cents per task. OpenAI recently slashed API pricing for GPT-5.6-Sol by twenty percent over a three-month period to aggressively compete with Anthropic. Muse Spark 1.2 is placed in the B tier, offering incredible value and blistering output speeds for specific coding and metadata tasks. Opus 5 drops down to the F tier because its high hallucination rate, lack of practical coding utility, and expensive pricing make it unusable for developers.
Context: Ben and Theo host Nerdsnipe, a developer-focused podcast analyzing frontier artificial intelligence models. In this episode, the two clash over a comprehensive model tier list, grading everything from OpenAI and Anthropic flagships to open weights and Chinese labs based on grueling real-world developer benchmarks.