OpenAI's GPT 5.4 in 10 Minutes: 1M Context, Computer Use, Coding Gains, Benchmarks & Pricing

Quick Overview

OpenAI released GPT-4o, featuring two new models, GPT-4o and GPT-4o-turbo, with GPT-4o available in ChatGPT Plus, Teams, and Enterprise, while GPT-4o-turbo requires the $200/month tier; key advancements include support for up to one million tokens of context, state-of-the-art computer use, and significant gains in reasoning and coding, evidenced by surpassing human performance on the OS World benchmark at 75%.

Key Points: GPT-4o supports up to one million tokens of context, and requests exceeding 272,000 tokens are charged at 2x the normal rate for those excess tokens. GPT-4o surpassed human performance on the OS World benchmark, achieving 75% verification compared to the human baseline of 72.4%. The model allows users to adjust its thinking mid-course while it is working, providing an upfront plan before arriving at a final output, which is described as an interesting UX aspect. GPT-4o-turbo is the most token-efficient model yet, which means it can potentially be cheaper despite having a higher base cost if it performs tasks with fewer tokens. API pricing for GPT-4o-turbo is high: $180 per million tokens of input and $180 per million tokens of output, compared to Claude Opus 4.6's $5 input and $25 output per million tokens. GPT-4o combines the coding strengths of GPT-4o-turbo with leading knowledge work and computer use capabilities, showing considerable leaps over GPT-4o-turbo in benchmarks, especially at the medium reasoning level. Lee Rob from Cursor stated that GPT-4o-turbo is the current leader on their internal benchmarks, noting it is more natural and assertive than previous models and proactively paralyses work.

Context: The video analyzes the newly released GPT-4o models from OpenAI, detailing their capabilities, availability tiers, and pricing structures in comparison to existing and competing models like Anthropic's Claude Opus 4.6. The speaker focuses heavily on advancements in context window size, reasoning, computer use, and coding performance, drawing direct comparisons using benchmark data and user feedback.

Raw markdown version of this recap