# The "Token Muncher" Problem: Is Sonnet 4.6 Actually Cheaper?

Source: https://www.youtube.com/watch?v=iyLwNnU6BMA
Recap page: https://rapidrecap.app/video/iyLwNnU6BMA
Generated: 2026-02-18T15:35:03.311+00:00

---
## Quick Overview

Claude Sonnet 4.6 is not definitively cheaper than Opus 4.6 for all tasks, as its token usage is significantly higher (280M tokens vs. 160M tokens for Opus 4.6 at equivalent performance levels on the GDPval-AA benchmark), making the overall cost potentially higher despite lower per-token pricing for some operations.

**Key Points:**
- Claude Sonnet 4.6 achieved an ELO of 1633 on the GDPval-AA benchmark using adaptive thinking mode, placing it slightly ahead of Anthropic's Opus 4.6 on agentic performance.
- To achieve this performance, Sonnet 4.6 used 280M tokens, which is more than 4x the tokens used by its predecessor, Sonnet 4.5 (58M tokens).
- Opus 4.6 achieved equivalent performance using only 160M tokens, meaning Sonnet 4.6 used approximately 40% more tokens than Opus 4.6 for comparable results.
- The token usage pushes Sonnet 4.6's total cost just ahead of Opus 4.6, despite Sonnet 4.6 having cheaper per-token pricing, because of the increased token consumption.
- The video suggests users should check the token usage charts, showing that models like GPT-5.2 Flash (Feb 2026) used the most tokens (320M) for the evaluation.
- The demonstration showed Sonnet 4.6 successfully handling agent tasks like updating delivery pricing and email elements using its tool-use capabilities.

![Screenshot at 00:10: The video displays a comparison table showing Sonnet 4.6 scoring 72.5% on the Agentic computer use \(OSWorld Verified\) benchmark, while Opus 4.6 scores 66.3%, illustrating the performance differences being discussed.](https://ss.rapidrecap.app/screens/iyLwNnU6BMA/00-00-10.jpg)

**Context:** The video discusses the release and performance of Anthropic's Claude Sonnet 4.6, contrasting it with the more powerful Opus 4.6 model, particularly focusing on the trade-off between per-token cost and token usage efficiency when performing complex agentic tasks. The speaker references benchmark results from Artificial Analysis's GDPval-AA leaderboard, which tracks model performance against token consumption.

## Detailed Analysis

The central theme of the video is evaluating whether the new Claude Sonnet 4.6 is truly cheaper than the higher-tier Opus 4.6, despite its lower per-token pricing. The speaker references external benchmark data (Artificial Analysis GDPval-AA) showing that while Sonnet 4.6 is the new leader on agentic performance (scoring ELO 1633, slightly ahead of Opus 4.6), its efficiency is lower. Specifically, Sonnet 4.6 used 280M tokens to achieve performance that Opus 4.6 reached with only 160M tokens, meaning Sonnet 4.6 consumed about 40% more tokens for similar results. Because token usage heavily influences total cost, the speaker concludes that Sonnet 4.6's total cost for these tasks pushes it just ahead of Opus 4.6, negating the expected cost savings from its cheaper per-token rate. The video also features a demonstration where Sonnet 4.6 successfully completes several administrative tasks, such as updating shipping rules and modifying email templates, showcasing its improved tool-use capabilities. The speaker advises users to consider the token usage implications, noting that while Sonnet 4.6 is excellent, Opus 4.6 might still be more cost-effective for complex, multi-step workflows due to its superior token efficiency.

### Performance Benchmarks

- Sonnet 4.6 leads Opus 4.6 on agentic performance (ELO 1633) but uses significantly more tokens (280M vs 160M).
- Sonnet 4.6 achieved a substantial improvement over Sonnet 4.5, using over 4x the tokens of its predecessor.

### Token Efficiency Comparison

- Opus 4.6 required 160M tokens for performance equivalent to Sonnet 4.6's 280M tokens, resulting in a ~40% less token usage for Opus 4.6.

### Pricing Analysis

- Despite Sonnet 4.6 having cheaper per-token rates, its higher token consumption means the total cost for complex tasks is slightly higher than Opus 4.6.

### Agentic Task Demonstration

- Sonnet 4.6 successfully processed a to-do list, including updating web store delivery pricing and modifying email header colors using programmatic tool calling.

### Future Outlook

- The speaker suggests users should still consider Opus 4.6 for tasks requiring long chains of thought, as the token cost difference might make Sonnet 4.6 less economical for those specific workloads.

![Screenshot at 00:04: Demonstration of Claude agent successfully executing a task by requesting to modify web store delivery pricing.](https://ss.rapidrecap.app/screens/iyLwNnU6BMA/00-00-04.jpg)
![Screenshot at 00:09: The agent navigates the StoreDesk dashboard, highlighting the 'Shipping & Delivery' section to address the pricing task.](https://ss.rapidrecap.app/screens/iyLwNnU6BMA/00-00-09.jpg)
![Screenshot at 00:23: The announcement slide for Claude Opus 4.6, comparing it visually against other models on a scatter plot.](https://ss.rapidrecap.app/screens/iyLwNnU6BMA/00-00-23.jpg)
![Screenshot at 01:01: A chart showing the steady improvement in Claude Sonnet's OSWorld score over time, culminating in Sonnet 4.6 reaching 72.5%.](https://ss.rapidrecap.app/screens/iyLwNnU6BMA/00-01-01.jpg)
![Screenshot at 02:12: The announcement page for Cowork: Claude Code, highlighting new features like file access and folder instructions for agents.](https://ss.rapidrecap.app/screens/iyLwNnU6BMA/00-02-12.jpg)
