GPT 5.4 "we see no wall"

Quick Overview

OpenAI released GPT-5.4, which shows significant performance gains across benchmarks like GDPVal, where it achieved an 83.0% win/tie rate on knowledge work tasks, surpassing human experts (49.8% baseline), and introduced native computer-use capabilities, achieving a 75.0% success rate on the OSWorld-Verified benchmark, exceeding GPT-5.2's 47.3%.

Key Points: GPT-5.4 was released on March 5, 2026, designed for professional work and available in ChatGPT, the API, and Codex. In GDPVal knowledge work tasks, GPT-5.4 achieved an 83.0% win/tie rate, significantly outperforming the industry expert baseline of 49.8%. GPT-5.4 features native computer-use capabilities, achieving a 75.0% success rate on the OSWorld-Verified benchmark, surpassing human performance (72.4%). The model also showed strong browser use performance on WebArena-Verified (67.3% success) and Online-Mind2Web (92.8% success). The release coincided with Anthropic's CEO statement challenging a Department of War designation regarding Claude as a supply chain risk. OpenAI also launched GPT-5.4 Thinking, featuring improved deep web research and the ability to interrupt and steer the model mid-response. Ryan Brewer announced joining OpenAI to build the future of financial intelligence using models like GPT-5.4 for financial reasoning and Excel-based modeling.

Context: The video discusses the release of OpenAI's new flagship model, GPT-5.4, on March 5, 2026, alongside related news from the AI industry. Key announcements include GPT-5.4's superior performance on productivity benchmarks (GDPVal), new computer vision/use capabilities, and the release of GPT-5.4 Thinking with steering controls. The context is further enriched by competitor news, specifically Anthropic challenging a Department of War designation and OpenAI launching financial service tools, indicating rapid, competitive advancement in the frontier AI space.

Detailed Analysis

The video covers the March 5, 2026 release of OpenAI's GPT-5.4, touted as their most capable and efficient frontier model for professional work, available in ChatGPT, API, and Codex. The speaker highlights performance metrics using the GDPVal benchmark, where GPT-5.4 achieved an 83.0% win/tie rate on knowledge work tasks, far exceeding the human expert baseline of 49.8% (GPT-5.4 Pro scored 82.0%). A major new feature is native computer-use capabilities, allowing agents to interact with websites and software. On the OSWorld-Verified benchmark, GPT-5.4 achieved a 75.0% success rate, surpassing human performance at 72.4% and GPT-5.2's 47.3%. Browser use benchmarks like WebArena-Verified (67.3%) and Online-Mind2Web (92.8%) also showed strong results. The speaker also mentions the concurrent release of GPT-5.4 Thinking, which includes better context retention and the ability to interrupt and steer the model mid-response. Furthermore, the video touches upon industry moves, including Anthropic CEO Dario Amodei's statement challenging the Department of War's supply chain risk designation for Claude, and the hiring of Ryan Brewer by OpenAI to build financial intelligence tools using GPT-5.4, indicating a focus on finance applications. Finally, the speaker touches upon safety research from Wes Roth detailing that advanced models struggle to control their internal 'thoughts' via Chain-of-Thought (CoT) controllability, which is seen as a positive safety feature.

Raw markdown version of this recap