# OpenAI's GPT 5.4 in 10 Minutes: 1M Context, Computer Use, Coding Gains, Benchmarks & Pricing

Source: https://www.youtube.com/watch?v=MwATr76kFXs
Recap page: https://rapidrecap.app/video/MwATr76kFXs
Generated: 2026-03-06T01:37:45.289+00:00

---
## Quick Overview

OpenAI released GPT-4o, featuring two new models, GPT-4o and GPT-4o-turbo, with GPT-4o available in ChatGPT Plus, Teams, and Enterprise, while GPT-4o-turbo requires the $200/month tier; key advancements include support for up to one million tokens of context, state-of-the-art computer use, and significant gains in reasoning and coding, evidenced by surpassing human performance on the OS World benchmark at 75%.

**Key Points:**
- GPT-4o supports up to one million tokens of context, and requests exceeding 272,000 tokens are charged at 2x the normal rate for those excess tokens.
- GPT-4o surpassed human performance on the OS World benchmark, achieving 75% verification compared to the human baseline of 72.4%.
- The model allows users to adjust its thinking mid-course while it is working, providing an upfront plan before arriving at a final output, which is described as an interesting UX aspect.
- GPT-4o-turbo is the most token-efficient model yet, which means it can potentially be cheaper despite having a higher base cost if it performs tasks with fewer tokens.
- API pricing for GPT-4o-turbo is high: $180 per million tokens of input and $180 per million tokens of output, compared to Claude Opus 4.6's $5 input and $25 output per million tokens.
- GPT-4o combines the coding strengths of GPT-4o-turbo with leading knowledge work and computer use capabilities, showing considerable leaps over GPT-4o-turbo in benchmarks, especially at the medium reasoning level.
- Lee Rob from Cursor stated that GPT-4o-turbo is the current leader on their internal benchmarks, noting it is more natural and assertive than previous models and proactively paralyses work.

**Context:** The video analyzes the newly released GPT-4o models from OpenAI, detailing their capabilities, availability tiers, and pricing structures in comparison to existing and competing models like Anthropic's Claude Opus 4.6. The speaker focuses heavily on advancements in context window size, reasoning, computer use, and coding performance, drawing direct comparisons using benchmark data and user feedback.

## Detailed Analysis

OpenAI launched two new models: GPT-4o and GPT-4o-turbo; GPT-4o is accessible in ChatGPT Plus, Teams, and Enterprise, while GPT-4o-turbo is restricted to the $200 per month tier, though both are available via API. Major features include a massive one million token context window and superior computer use capabilities, allowing the model to provide an upfront thinking plan that users can adjust mid-course without needing additional turns. Benchmarks show GPT-4o is an incremental improvement over GPT-4o-turbo codecs, notably surpassing human performance on OS World at 75%. The model excels in real-world knowledge work, generating more consistent and polished documents and presentations than previous versions. For coding, GPT-4o integrates the strengths of GPT-4o-turbo, showing significant improvements, including the ability to generate complex front-end applications like theme park simulations and RPG games. Pricing is steep for the top-tier model: $180 per million tokens for both input and output, significantly more expensive than Claude Opus 4.6 ($5 input/$25 output). Despite the cost, early reactions, including from Lee Rob at Cursor, are overwhelmingly positive, confirming its leadership in internal benchmarks.

### Model Availability and Tiers

- GPT-4o is available in Chat GPT plus, teams, pro, and enterprise
- GPT-4o-turbo is exclusively available in the $200 a month tier
- Both models are accessible from the API

### Key Technical Advancements

- Supports up to one million tokens of context
- State-of-the-art computer use capabilities
- Improved tool search capability and MCP support within cloud code

### User Interaction and UX

- Model provides an upfront plan of its thinking
- Users can adjust mid-course while the model is working to align the final output
- This is compared to Anthropic's interleaf thinking capability

### Benchmark Performance

- Surpassed humans on OS World benchmark at 75% verification
- Model is quite a bit better across browse comp and web arena tasks
- Shows considerable leap over GPT-4o-turbo on reasoning tasks, shining on medium reasoning level

### Knowledge Work and Generation Quality

- Delivers more consistent and polished results on real world work
- Generated documents and presentations look much more professional and cohesive than previous models
- Demonstrates ability to build complex code like theme park simulations and RPG games

### API Pricing Structure

- GPT-4o-turbo costs $180 per million tokens of output and $180 per million tokens of input
- Input requests exceeding 272,000 tokens of context are charged at 2x the normal rate
- Claude Opus 4.6 is significantly cheaper at $5 per million tokens of input and $25 per million tokens of output

### Developer Feedback

- Lee Rob from cursor confirmed GPT-4o-turbo is the leader on internal benchmarks
- Engineers find it more natural and assertive than previous models
- It works through ambiguous problems without second-guessing itself

