GPT-5 is here... Can it win back programmers?

Quick Overview

GPT-5 is not the groundbreaking leap forward many expected, failing to achieve human-level performance in key benchmarks and exhibiting significant flaws in its programming assistance capabilities, leading to questions about its value and whether it's merely an overhyped incremental update.

Key Points: GPT-5's "thinking" capability is highlighted as a key differentiator, but its performance on benchmarks like SWE-bench Verified (Software Engineering) is shown to be lower than expected, scoring 52.8% "with thinking" compared to a human baseline of 83.7%. The model's performance on the ARC-AGI-2 benchmark is also questioned, with a claim that it performed poorly on a cost-per-task graph, scoring 9.9% compared to Grok 4's 15.9%. Concerns are raised about the accuracy and presentation of OpenAI's own benchmark charts, specifically an error in the Y-axis scale for the "Coding deception" metric, suggesting either a lack of PhD-level intelligence or intentional misrepresentation. Despite claims of being the "smartest model yet," GPT-5 struggled with basic programming tasks, failing to generate a functional Svelte app and producing a 500 internal error, needing manual correction. GPT-5 is priced at $10/M tokens, significantly more expensive than alternatives like Claude Opus 4.1 at $75/M tokens, raising questions about its cost-effectiveness given its performance limitations. The video suggests that parameter scaling, the traditional method for improving AI models, may be dead, and that GPT-5's advancements are more about consolidating and efficiently routing multiple models rather than a fundamental intelligence increase. The overall sentiment is that GPT-5, while an improvement, is not the revolutionary AI advancement it was hyped to be, failing to deliver on promises of human-level intelligence or significantly better programming assistance.

Context: This video critically examines OpenAI's GPT-5, questioning its claimed advancements and performance against competitors like Grok 4 and Claude Opus. It delves into benchmark results, pricing, and practical coding demonstrations to assess whether GPT-5 represents a true leap in AI capabilities or is simply an incremental upgrade with significant hype. The video highlights potential flaws in OpenAI's own data presentation and GPT-5's struggles with real-world coding tasks, contrasting these with its high price point and the purported benefits of its "thinking" feature.

Raw markdown version of this recap