# GPT 5.4 "we see no wall"

Source: https://www.youtube.com/watch?v=9zVZVtPMU6Y
Recap page: https://rapidrecap.app/video/9zVZVtPMU6Y
Generated: 2026-03-06T05:01:17.596+00:00

---
## Quick Overview

OpenAI released GPT-5.4, which shows significant performance gains across benchmarks like GDPVal, where it achieved an 83.0% win/tie rate on knowledge work tasks, surpassing human experts (49.8% baseline), and introduced native computer-use capabilities, achieving a 75.0% success rate on the OSWorld-Verified benchmark, exceeding GPT-5.2's 47.3%.

**Key Points:**
- GPT-5.4 was released on March 5, 2026, designed for professional work and available in ChatGPT, the API, and Codex.
- In GDPVal knowledge work tasks, GPT-5.4 achieved an 83.0% win/tie rate, significantly outperforming the industry expert baseline of 49.8%.
- GPT-5.4 features native computer-use capabilities, achieving a 75.0% success rate on the OSWorld-Verified benchmark, surpassing human performance (72.4%).
- The model also showed strong browser use performance on WebArena-Verified (67.3% success) and Online-Mind2Web (92.8% success).
- The release coincided with Anthropic's CEO statement challenging a Department of War designation regarding Claude as a supply chain risk.
- OpenAI also launched GPT-5.4 Thinking, featuring improved deep web research and the ability to interrupt and steer the model mid-response.
- Ryan Brewer announced joining OpenAI to build the future of financial intelligence using models like GPT-5.4 for financial reasoning and Excel-based modeling.

![Screenshot at 00:00: The title slide announcing "Introducing GPT-5.4" designed for professional work, setting the stage for the model's release details.](https://ss.rapidrecap.app/screens/9zVZVtPMU6Y/00-00-00.jpg)

**Context:** The video discusses the release of OpenAI's new flagship model, GPT-5.4, on March 5, 2026, alongside related news from the AI industry. Key announcements include GPT-5.4's superior performance on productivity benchmarks (GDPVal), new computer vision/use capabilities, and the release of GPT-5.4 Thinking with steering controls. The context is further enriched by competitor news, specifically Anthropic challenging a Department of War designation and OpenAI launching financial service tools, indicating rapid, competitive advancement in the frontier AI space.

## Detailed Analysis

The video covers the March 5, 2026 release of OpenAI's GPT-5.4, touted as their most capable and efficient frontier model for professional work, available in ChatGPT, API, and Codex. The speaker highlights performance metrics using the GDPVal benchmark, where GPT-5.4 achieved an 83.0% win/tie rate on knowledge work tasks, far exceeding the human expert baseline of 49.8% (GPT-5.4 Pro scored 82.0%). A major new feature is native computer-use capabilities, allowing agents to interact with websites and software. On the OSWorld-Verified benchmark, GPT-5.4 achieved a 75.0% success rate, surpassing human performance at 72.4% and GPT-5.2's 47.3%. Browser use benchmarks like WebArena-Verified (67.3%) and Online-Mind2Web (92.8%) also showed strong results. The speaker also mentions the concurrent release of GPT-5.4 Thinking, which includes better context retention and the ability to interrupt and steer the model mid-response. Furthermore, the video touches upon industry moves, including Anthropic CEO Dario Amodei's statement challenging the Department of War's supply chain risk designation for Claude, and the hiring of Ryan Brewer by OpenAI to build financial intelligence tools using GPT-5.4, indicating a focus on finance applications. Finally, the speaker touches upon safety research from Wes Roth detailing that advanced models struggle to control their internal 'thoughts' via Chain-of-Thought (CoT) controllability, which is seen as a positive safety feature.

### GPT-5.4 Release & Performance

- GPT-5.4 launched on March 5, 2026, in ChatGPT, API, and Codex
- GDPVal knowledge work tasks showed 83.0% win/tie rate for GPT-5.4, beating the 49.8% human baseline
- GPT-5.4 Pro scored 82.0% in GDPVal.

### Computer Use & Vision Capabilities

- GPT-5.4 achieved 75.0% success on OSWorld-Verified (desktop navigation), surpassing human performance (72.4%)
- Excellent at writing code for computer operation via libraries like Playwright
- Demonstrated success in troubleshooting game development issues.

### New GPT-5.4 Thinking Features

- Features improved deep web research and better context retention
- Introduces 'Steering' allowing users to interrupt and adjust model direction mid-response.

### Industry Competitor News

- Anthropic CEO challenged the Department of War's supply chain risk designation for Claude
- Bloomberg reported OpenAI releasing financial services tools and a new flagship model.

### Finance and Benchmarks

- Ryan Brewer joined OpenAI to build financial intelligence tools using GPT-5.4 (Excel/financial reasoning)
- OpenAI's internal investment banking benchmark shows GPT-5.4 Thinking scoring 0.873, the highest among tested models.

### Safety Research Highlight

- Wes Roth shared research showing advanced models struggle to control their internal Chain-of-Thought (CoT) compared to output controlability, viewed as a positive safety measure.

![Screenshot at 00:04: The announcement slide for GPT-5.4, highlighting its release date \(March 5, 2026\) and focus on professional work.](https://ss.rapidrecap.app/screens/9zVZVtPMU6Y/00-00-04.jpg)
![Screenshot at 00:21: A bar chart from the GDPVal evaluation showing GPT-5.4 Pro achieving an 82.0% win rate against an industry expert baseline of 49.8% on knowledge work tasks.](https://ss.rapidrecap.app/screens/9zVZVtPMU6Y/00-00-21.jpg)
![Screenshot at 04:33: Text describing GPT-5.4's native computer-use capabilities and its 75.0% success rate on the OSWorld-Verified benchmark.](https://ss.rapidrecap.app/screens/9zVZVtPMU6Y/00-04-33.jpg)
![Screenshot at 08:06: A tweet from Bloomberg announcing OpenAI's new flagship model and a suite of financial-services tools.](https://ss.rapidrecap.app/screens/9zVZVtPMU6Y/00-08-06.jpg)
