Claude Opus 4.6 vs GPT 5.3 Codex: Which is better for programming? | Peter Steinberger

Quick Overview

The discussion concludes that Claude Opus 4.6 is generally superior to GPT-5.3-Codex for programming tasks, particularly because Opus is more pleasant to use and requires less explicit guidance, although Codex is capable of producing good results when heavily prompted and is more readily available in certain environments.

Key Points: Claude Opus 4.6 is considered more pleasant to use than GPT-5.3-Codex for programming due to its more natural interaction style. The speaker notes that Codex requires more explicit guidance and effort to achieve good results, whereas Opus is more intuitive. While Codex might be slightly slower, the quality difference suggests that if you pay for the $200 version of Codex, you get the slower version, while the cheaper open version is faster but less capable. The discussion touches upon the psychological aspect where users fall in love with a new model (like Codex) until they experience a superior alternative (like Opus). The speaker prefers the experience of Opus because it feels less like a trial-and-error process and more like a natural collaboration. The concept of 'quantifying' the intelligence of models like Codex is difficult, but the speaker prefers Opus's 'feel' and interactivity.

Context: This video features Lex Fridman interviewing Peter Steinberger about the comparative performance and user experience of two large language models for programming: Anthropic's Claude Opus 4.6 and OpenAI's GPT-5.3-Codex. The conversation centers on which model offers a better developer experience, focusing on factors like code quality, required prompting effort, and overall interaction 'feel,' drawing comparisons to previous models and the effort required to extract quality output.

Detailed Analysis

The discussion centers on comparing Claude Opus 4.6 and GPT-5.3-Codex for programming. Peter Steinberger strongly favors Opus 4.6, stating it is more 'pleasant to use' because it requires less explicit steering and feels more like a natural co-pilot. He contrasts this with Codex, which, while capable, often requires significantly more effort, specific prompting, and trial-and-error to achieve good results. Steinberger notes that if users pay for the premium version of Codex, they might experience slower performance compared to the cheaper, faster, but less capable open version. He also mentions the psychological effect where people initially love a new model (like Codex) until a better one (like Opus) emerges, leading to user dissatisfaction with the older model. He suggests that the intelligence difference between the models is hard to quantify precisely, but the interactive experience and 'feel' heavily favor Opus, which is less likely to degrade its performance over time when dealing with complex tasks or refactoring.

Raw markdown version of this recap