# ChatGPT vs Gemini vs Claude vs Grok: Who wins? | Lex Fridman Podcast

Source: https://www.youtube.com/watch?v=iD6IYG8W5jY
Recap page: https://rapidrecap.app/video/iD6IYG8W5jY
Generated: 2026-02-03T03:34:03.918+00:00

---
## Quick Overview

The discussion concludes that while GPT-5.2 currently leads in a simulated coding benchmark (LMArena Code) with a 1475 score, Google Gemini is favored by the guest for its superior performance in long-context retrieval tasks like the Needle in a Haystack test, suggesting Gemini's approach may be better for certain complex, real-world applications despite OpenAI's perceived lead in raw benchmarks.

**Key Points:**
- GPT-5.2 scored 1475 on the LMArena Code benchmark, ranking second only to Claude 3 Opus (1504) as of the data presented.
- The guest personally prefers Gemini for its performance in long-context retrieval tasks, like the Needle in a Haystack test, where they observed superior consistency.
- The guest noted that GPT-5.2's 'Thinking' mode shows high accuracy (near 100% at smaller context sizes) in the Needle in a Haystack test, but performance drops significantly as context length increases to 256k tokens.
- Google's TPU infrastructure (like the Ironwood superpod with 9,216 TPUs) gives them a potential hardware advantage for training large models efficiently, challenging NVIDIA's GPU dominance.
- The speaker mentioned that for personal development work, they often use Gemini for fast tasks (Instant mode) but switch to GPT-5.2's 'Thinking' or 'Pro' mode for more complex, error-prone tasks.
- The conversation touched upon the different philosophies: OpenAI focusing on raw intelligence/speed (GPT-5), and Google focusing on scaling infrastructure (TPUs) to maintain a competitive edge.

![Screenshot at 07:59: A slide comparing coding benchmark scores shows GPT-5.2 at rank 2 \(1475 score\) behind Claude 3 Opus, illustrating the competitive state of top models in code generation tasks.](https://ss.rapidrecap.app/screens/iD6IYG8W5jY/00-07-59.jpg)

**Context:** This segment features Lex Fridman interviewing an AI researcher/developer, likely from DeepSeek-AI based on the initial slide, discussing the current state and competitive landscape of frontier Large Language Models (LLMs) as of early 2025 projections. The discussion centers around comparing models like OpenAI's GPT-5.2, Google's Gemini, Anthropic's Claude, and Grok, using specific benchmarks like LMArena Code and the 'Needle in a Haystack' test to evaluate capabilities in coding, reasoning, and long-context recall.

## Detailed Analysis

The discussion compares the perceived strengths of leading LLMs. Lex Fridman presents data suggesting GPT-5.2 is highly competitive, ranking second on the LMArena Code benchmark with a score of 1475, only slightly behind Claude 3 Opus (1504). However, the guest argues that for practical application, especially involving long contexts, Gemini shows advantages. Specifically regarding the 'Needle in a Haystack' test, which measures retrieval accuracy across long contexts, the guest points out that while GPT-5.2's 'Thinking' mode performs nearly perfectly at smaller context windows (e.g., 8k tokens), its performance degrades more sharply than competitors as context length extends to 256k tokens. The guest admits to heavily favoring Gemini for tasks requiring deep context retention, even though they use GPT-5.2's faster modes for simpler daily tasks. Furthermore, the conversation highlights Google's infrastructural advantage with their custom TPUs (like the Ironwood superpod), which allows them to scale training more cost-effectively than relying solely on external hardware like NVIDIA GPUs, potentially enabling them to drive innovation in areas like multi-modality and long-context handling.

### LLM Benchmarking & User Preference

- GPT-5.2 scores 1475 on LMArena Code, ranking second
- Guest prefers Gemini for long-context retrieval (Needle in a Haystack) due to better consistency across large context sizes
- Guest uses GPT-5.2's 'Thinking' mode for complex coding tasks and Gemini for fast tasks, though they sometimes default to GPT-5.2.

### Context Window Performance

- GPT-5.2 Thinking mode shows near 100% accuracy at 8k tokens but drops significantly by 256k tokens in the MRCv2 test
- This degradation suggests trade-offs between speed/simplicity and deep context retention.

### Infrastructure & Competition

- Google's TPU architecture, exemplified by the Ironwood superpod, allows them to train frontier models without relying on NVIDIA GPUs, potentially leading to better cost margins and faster iteration on proprietary hardware.

### Model Philosophy

- The difference in performance suggests a trade-off between raw intelligence (where US models seem stronger) and the ability to scale efficiently (where Google's infrastructure is highlighted).

![Screenshot at 00:03: Title slide introducing DeepSeek-V3.2 and outlining key technical breakthroughs like Sparse Attention and Scalable Reinforcement Learning Framework.](https://ss.rapidrecap.app/screens/iD6IYG8W5jY/00-00-03.jpg)
![Screenshot at 00:19: A bar chart projecting AI Chatbot Market Share Worldwide for January 2026, showing ChatGPT dominating at 80.14%, Perplexity at 8.15%, and DeepSeek at 0.01%.](https://ss.rapidrecap.app/screens/iD6IYG8W5jY/00-00-19.jpg)
![Screenshot at 03:57: A table detailing the Google TPU Architecture Evolution and Scaling Timeline, showing performance gains from v1 \(2015\) up to v7 Ironwood \(2025\).](https://ss.rapidrecap.app/screens/iD6IYG8W5jY/00-03-57.jpg)
![Screenshot at 03:48: A line graph comparing the TFLOPs and System Availability of Google TPUs versus NVIDIA GPUs from Jan 2020 to Jan 2026, showing NVIDIA's H100/GB200 leading NVIDIA GPU performance over TPUs, but TPUs showing aggressive recent growth with v7.](https://ss.rapidrecap.app/screens/iD6IYG8W5jY/00-03-48.jpg)
![Screenshot at 04:38: A slide from OpenAI introducing GPT-5, detailing it as a unified system with a smart, efficient model, a deeper reasoning model, and a real-time router.](https://ss.rapidrecap.app/screens/iD6IYG8W5jY/00-04-38.jpg)
