My GPT-5 First Reaction - It's smarter but there's a few things missing...
Quick Overview
David Shapiro reviews the GPT-5 launch stream, finding it underwhelming in multimodality and agentic behavior, but praising its reduced hallucination and improved reasoning, though he notes a lack of focus on practical applications and math benchmarks.
Key Points: GPT-5 shows impressive progress in reducing hallucination and improving reasoning capabilities. The launch presentation lacked focus on multimodality, agentic behavior, and practical applications. Performance benchmarks indicate exponential improvement in LLM task completion times. Shapiro suggests a need for a clearer normative framework and consideration of broader societal impacts. The model's ability to handle larger context windows and complex tasks is a significant advancement. The presentation was perceived as overly focused on 'normies' and lacking technical depth. Shapiro questions the completeness of the current framework and calls for consideration of human rights and environmental factors.
Context: David Shapiro, a commentator on technology and economics, shares his thoughts on the recent GPT-5 launch livestream. He analyzes the presentation, the model's capabilities, and the broader implications for AI development, drawing comparisons to previous models and existing frameworks.
Detailed Analysis
David Shapiro shares his initial reaction to the GPT-5 launch stream, expressing disappointment with the lack of emphasis on multimodality and agentic behavior, noting that these were not front and center or even major advancements. He also found the agentic behavior aspect to be underdeveloped, with only one graph showing how much more agentic tasks could be performed. While he acknowledges the rise in benchmarks and coding capability, he feels it's only slightly better than expected and reserves final judgment until he can get his hands on it. Shapiro believes the presentation was "watered down" and felt more like a PR event for "normies" than a deep dive into technical advancements. However, he is "really excited" about the lower levels of hallucination and sycophancy, combined with higher intelligence, better instruction following, and larger context windows, which he believes will have significant "knock-on effects" for project size. He also points out the omission of math benchmarks and the lack of demonstration for practical applications like coding or reasoning. Shapiro highlights a "METR" chart showing the time-horizon of software engineering tasks different LLMs can complete, noting the exponential progress from GPT-2 to GPT-5, with GPT-5 performing tasks in minutes that GPT-2 took hours for. He also critiques the presentation for not fully exploring the framework's potential or limitations, suggesting that a more robust approach would involve a multi-equilibrium view and a clearer "north star" for development. Shapiro believes that while the progress is impressive, there's a missing normative anchor and a lack of focus on real-world applications beyond benchmarks.