# GPT-5 is here... Can it win back programmers?

Source: https://www.youtube.com/watch?v=8tx2viHpgA8
Recap page: https://rapidrecap.app/video/8tx2viHpgA8
Generated: 2025-08-28T10:29:43.743+00:00

---
## Quick Overview

GPT-5 is not the groundbreaking leap forward many expected, failing to achieve human-level performance in key benchmarks and exhibiting significant flaws in its programming assistance capabilities, leading to questions about its value and whether it's merely an overhyped incremental update.

**Key Points:**
- GPT-5's "thinking" capability is highlighted as a key differentiator, but its performance on benchmarks like SWE-bench Verified (Software Engineering) is shown to be lower than expected, scoring 52.8% "with thinking" compared to a human baseline of 83.7%.
- The model's performance on the ARC-AGI-2 benchmark is also questioned, with a claim that it performed poorly on a cost-per-task graph, scoring 9.9% compared to Grok 4's 15.9%.
- Concerns are raised about the accuracy and presentation of OpenAI's own benchmark charts, specifically an error in the Y-axis scale for the "Coding deception" metric, suggesting either a lack of PhD-level intelligence or intentional misrepresentation.
- Despite claims of being the "smartest model yet," GPT-5 struggled with basic programming tasks, failing to generate a functional Svelte app and producing a 500 internal error, needing manual correction.
- GPT-5 is priced at $10/M tokens, significantly more expensive than alternatives like Claude Opus 4.1 at $75/M tokens, raising questions about its cost-effectiveness given its performance limitations.
- The video suggests that parameter scaling, the traditional method for improving AI models, may be dead, and that GPT-5's advancements are more about consolidating and efficiently routing multiple models rather than a fundamental intelligence increase.
- The overall sentiment is that GPT-5, while an improvement, is not the revolutionary AI advancement it was hyped to be, failing to deliver on promises of human-level intelligence or significantly better programming assistance.

![Screenshot at 00:15: A bar graph titled "SimpleBench: Model Scores vs Human Baseline" shows GPT-5 \(Copilot\) scoring 90%, significantly higher than the human baseline at 83.7%, and much higher than other AI models like Gemini 2.5 Pro and Claude 4 Opus.](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-00-15.png)

**Context:** This video critically examines OpenAI's GPT-5, questioning its claimed advancements and performance against competitors like Grok 4 and Claude Opus. It delves into benchmark results, pricing, and practical coding demonstrations to assess whether GPT-5 represents a true leap in AI capabilities or is simply an incremental upgrade with significant hype. The video highlights potential flaws in OpenAI's own data presentation and GPT-5's struggles with real-world coding tasks, contrasting these with its high price point and the purported benefits of its "thinking" feature.

## Detailed Analysis

The video critically analyzes the release of OpenAI's GPT-5, challenging the assertion that it represents a significant leap in AI capabilities. Initial claims positioned GPT-5 as the "smartest, fastest, and most useful model yet," boasting "thinking built-in." However, the video presents evidence suggesting these claims are largely unsubstantiated. Performance benchmarks, such as the SWE-bench Verified test for software engineering, show GPT-5 scoring 52.8% with thinking, falling considerably short of the human baseline at 83.7% and even behind other models. Further doubts are cast on its performance in other benchmarks, with a comparison showing Grok 4 outperforming GPT-5 on a cost-per-task metric. The video also points out apparent inaccuracies in OpenAI's own benchmark charts, specifically a mislabeled Y-axis on a "Coding deception" graph, which raises questions about the integrity of the data presented. In practical demonstrations, GPT-5 struggled to generate a functional Svelte app, resulting in errors that required manual correction, contrary to its advertised programming assistance capabilities. The pricing of GPT-5 at $10/M tokens is also highlighted as being substantially higher than competitors like Claude Opus 4.1, which costs $75/M tokens, making its performance-to-cost ratio questionable. The analysis suggests that the advancement in GPT-5 is less about genuine intelligence increase and more about improved model routing and consolidation, possibly indicating that traditional parameter scaling is reaching its limits. Ultimately, the video concludes that GPT-5, while an improvement, does not live up to the extensive hype and may not be the game-changer many anticipated, especially for programmers.

### Performance Benchmarks

- GPT-5 scores 52.8% on SWE-bench (with thinking), significantly below human baseline (83.7%)
- Grok 4 outperforms GPT-5 on ARC-AGI-2 cost-per-task benchmark (15.9% vs 9.9%)

### Data Integrity Concerns

- OpenAI's own benchmark charts show potential errors, like a mislabeled Y-axis on a 'Coding deception' metric, raising questions about accuracy and intentional misrepresentation

### Practical Coding Tests

- GPT-5 fails to generate functional Svelte app code, producing errors and requiring manual fixes, contradicting claims of superior programming assistance

### Cost Analysis

- GPT-5 is priced at $10/M tokens, significantly higher than alternatives like Claude Opus 4.1 at $75/M tokens

### Model Architecture & Strategy

- Parameter scaling may be obsolete; GPT-5 focuses on model routing and consolidation rather than pure intelligence increase

### Overall Assessment

- GPT-5 is an incremental update, not a revolutionary AI leap, failing to meet hype and struggling with practical applications despite high cost and claims of "thinking" capability

![Screenshot at 00:15: A bar graph titled "SimpleBench: Model Scores vs Human Baseline" shows GPT-5 \(Copilot\) scoring 90%, significantly higher than the human baseline at 83.7%, and much higher than other AI models like Gemini 2.5 Pro and Claude 4 Opus.](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-00-15.png)
![Screenshot at 00:20: A leaderboard overview displays GPT-5 ranking first in both 'Text' and 'WebDev' categories with scores of 1481 and 1479 respectively, ahead of Gemini 2.5 Pro.](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-00-20.png)
![Screenshot at 00:34: A leaderboard table shows GPT-5 \(high\) ranked 5th with a score of 56.7%, below the human baseline \(83.7%\) and several other models like Gemini 2.5 Pro and Grok 4.](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-00-34.png)
![Screenshot at 00:38: A tweet displays an "ARC-AGI-2 LEADERBOARD" showing Grok 4 outperforming GPT-5 on a score vs. cost-per-task graph.](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-00-38.png)
![Screenshot at 00:45: A Polymarket graph shows the market prediction for "Which company has best AI model end of 2025?", with OpenAI's prediction dropping significantly in July.](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-00-45.png)
![Screenshot at 02:03: A bar chart labeled "SWE-bench Verified" shows GPT-5 scoring 52.8% with thinking, while OpenAI o3 scores 69.1% and GPT-4o scores 30.8%.](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-02-03.png)
![Screenshot at 02:17: A bar chart shows "Coding deception" rates, with GPT-5 at 50.0% and another model at 47.4%, accompanied by text "C'mon Bro".](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-02-17.png)
![Screenshot at 02:38: The ChatGPT 5 interface shows a prompt to "Ask anything", with the text "build a svelte 5 app" being typed into the input field.](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-02-38.png)
![Screenshot at 02:44: A code editor displays a Svelte component with an error message in the terminal: "Error when evaluating SSR module... '$derived\(...\)' can only be used as a variable declaration initializer..."](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-02-44.png)
![Screenshot at 03:11: A functional To-Do app built with Svelte displays three items: 'learn to read good', 'build bunker', and 'make video'. The "make video" item is checked, reducing the count to 2.](https://ss.rapidrecap.app/screens/8tx2viHpgA8/00-03-11.png)
