# Mar 5 GPT 5.4 Drops LIVE on the Show! + Anthropic vs Pentagon, Qwen Drama, Gemini 3.1 Flash-Lite

Source: https://www.youtube.com/watch?v=SVimfToLUjA
Recap page: https://rapidrecap.app/video/SVimfToLUjA
Generated: 2026-03-06T02:03:30.388+00:00

---
## Quick Overview

OpenAI launched the new generalized model GPT 5.4 Thinking and 5.4 Pro, which incorporates the coding capabilities of GPT 5.3 Codec and features a 1 million token context window, while simultaneously, the Anthropic vs. Pentagon dispute escalated with OpenAI securing a deal to deploy its models instead of Claude, leading to community backlash.

**Key Points:**
- OpenAI dropped GPT 5.4 Thinking and 5.4 Pro, folding in the coding capabilities of GPT 5.3 Codec into a unified reasoning model, which reportedly solved a 20-year research-level math problem for Polish mathematician Bartosh Nazerki.
- The GPT 5.4 Pro version hit 83.3% on MMLU RKGI 2, closely matching Gemini Deep Think, and scored 75% on OS World and computer use, supposedly beating the human baseline.
- GPT 5.4 features a 1 million context window, costs $2.50 for input tokens, and introduces 'mid-thought steering' on the ChatGPT interface, allowing real-time redirection of the model's reasoning.
- Anthropic refused the Pentagon's requests regarding spying on citizens and using Claude in autonomous weapon kill chains, leading Secretary of Defense Pete Hex to designate Anthropic as a supply chain risk.
- Following Anthropic's refusal, Sam Altman announced OpenAI secured a deal with the Department of War, leading to immediate backlash with trending hashtags like #quitOpenAI, forcing OpenAI to amend the deal on Monday to add surveillance and weapons prohibitions.
- Alibaba's Qwen lab released the Qwen 3.5 small model series, with the 9B parameter model natively multimodal and rivaling GPT-4 120B, immediately followed by the departure of its tech lead, Jun Yang, which Alibaba's CEO is now addressing by co-leading the lab.
- Wolf from introduced 'Wolfbench,' an evaluation framework based on Terminal Bench, revealing that while Claude Opus 4.6 excels in maximum task solving (88% ceiling), its reliability (tasks solved 100% of the time) is only 55%, indicating consistency issues across different harness environments.

**Context:** The video is a news recap show hosted by Alex Vulov with co-hosts LDJ and Wolf from, covering major developments in the AI industry as of March 5th, focusing heavily on new model releases, corporate drama involving national security implications, and open-source advancements. Key events include the release of OpenAI's GPT 5.4, the highly publicized conflict between Anthropic and the US Department of Defense, and significant leadership changes at Alibaba's Qwen lab.

## Detailed Analysis

The primary news was the live drop of OpenAI's GPT 5.4 Thinking and 5.4 Pro, a unified model that integrates GPT 5.3 Codec capabilities, boasting a 1 million token context window and introducing real-time reasoning interruption via mid-thought steering in the ChatGPT interface. Performance metrics show GPT 5.4 Pro achieving 83.3% on MMLU RKGI 2 and 75% on computer use benchmarks. Concurrently, the Anthropic/Pentagon saga intensified after Anthropic refused DoD demands about surveillance and autonomous weapons; the DoD designated Anthropic a supply chain risk, prompting Sam Altman to announce OpenAI's deal to replace Anthropic, which caused massive public backlash against OpenAI until they amended the deal to include safety prohibitions. In open source, Alibaba released the highly capable, natively multimodal Qwen 3.5 small models, but this was overshadowed by the public resignation of Qwen tech lead Jun Yang, leading Alibaba's CEO to step in and co-lead the lab to assure continued open-source commitment. Finally, Wolf from presented 'Wolfbench,' a framework showing that while top models like Claude Opus perform best in ideal conditions on Terminal Bench, their consistency (tasks solved every run) lags significantly behind their maximum potential, highlighting reliability differences across different agent harnesses like CodeX and OpenCloud.

### GPT 5.4 Release Details

- Features 1 million context window
- Costs $2.50 input
- Introduced mid-thought steering on ChatGPT interface
- Backed by solving a 20-year frontier math problem

### Anthropic vs. Pentagon Conflict

- Anthropic refused DoD requests on citizen spying and autonomous kill chains
- DoD designated Anthropic a supply chain risk
- OpenAI secured a deal to deploy its models instead, facing immediate public backlash

### Qwen 3.5 Small Model Launch

- Released Qwen 3.5 small series with native multimodal capabilities
- 9B parameter model rivals GPT-4 120B on benchmarks
- Tech lead Jun Yang resigned immediately following the launch

### Google Gemini 3.1 Flashlight

- Launched as the fastest and most cost-efficient model in the 3.1 series
- Achieves 360 tokens per second
- Competes with Haiku and GPT-3.5 Mini on benchmarks, though pricing increased significantly over previous Flash versions

### Open Source and Agent Tools

- Qwen 3.5 models are highly usable on consumer GPUs (e.g., 15 tokens/sec for 27B on a 3090)
- Google released the Workspace CLI, signaling support for agentic tool use
- Symfony, an orchestration layer spec, was released for agents to build themselves

### Wolfbench Evaluation Framework

- Framework focuses on reliability and consistency across multiple runs, not just average scores
- Claude Opus 4.6 shows a high ceiling (88% max solve) but low consistency (55% always solved) on Terminal Bench
- Performance varies significantly depending on the harness used (Terminal Bench vs. CodeX vs. OpenCloud)

