# Deepseek just killed LLMs

Source: https://www.youtube.com/watch?v=4D-AsJ5UhF4
Recap page: https://rapidrecap.app/video/4D-AsJ5UhF4
Generated: 2025-10-23T08:32:33.184+00:00

---
## Quick Overview

The video analyzes recent developments in AI, focusing on DeepSeek-OCR's optical context compression for LLMs, Andrej Karpathy's critique of tokenizers and promotion of image-based inputs, Google AI's quantum computing breakthrough, and the controversy surrounding the 'Definition of AGI' paper's fabricated citations, all while highlighting the rapid progress and associated safety/infrastructure concerns in the AI field.

**Key Points:**
- DeepSeek-OCR achieves up to 20x compression of visual contexts while maintaining 97% OCR accuracy at context lengths under 10k tokens, demonstrating efficient optical context compression for LLMs.
- Andrej Karpathy champions image-based inputs over text tokens, arguing text tokenizers are wasteful and introduce historical baggage like security/jailbreak risks, suggesting pure image inputs are superior.
- Google AI announced a major quantum computing breakthrough, demonstrating a quantum computer running a verifiable algorithm 13,000x faster than leading classical supercomputers.
- The viral 'Definition of AGI' paper faced scrutiny after Michael Saxon pointed out it contained fake citations, which co-author Dan Hendrycks attributed to an incorrect conversion from a Google Doc to BibTeX.
- The video contrasts these advancements with infrastructure scaling, noting that training models like the C2S-Scale 27B costs as little as $100, but scaling laws suggest costs will increase quadratically with sequence length.
- The video also references the C2S-Scale 27B model's success in discovering a novel cancer therapy pathway, validating the power of scaling laws even in biological discovery.
- The discussion touches on the debate between token-based and pixel-based inputs, with Karpathy strongly favoring pixels due to tokenization's inherent issues like producing 'weird tokens' for emojis.

![Screenshot at 00:09: Speaker reacts to the DeepSeek-OCR announcement, illustrating the efficiency gains in optical context compression for LLMs.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-00-09.png)

**Context:** This video provides a commentary and digest of recent, high-profile news and research papers in the field of Artificial Intelligence, primarily focusing on Large Language Models (LLMs) and multimodal AI. Key subjects include the release of DeepSeek-OCR, Andrej Karpathy's ongoing thoughts on model input modalities (text vs. pixels), major announcements from Google AI regarding quantum computing, and a controversy surrounding a widely circulated AGI definition paper that contained fake citations. The speaker analyzes these topics by referencing the original tweets and research abstracts.

## Detailed Analysis

The video begins by discussing the DeepSeek-OCR paper, highlighting its success in compressing visual contexts up to 20x while retaining 97% OCR accuracy, which is crucial for efficient multimodal LLMs. The speaker then transitions to Andrej Karpathy's critique of text tokenization, referencing his tweet where he suggests that pixels are inherently better inputs than text tokens because tokenizers introduce issues like historical baggage, security risks, and inefficient representation (e.g., emojis becoming 'weird tokens'). Karpathy suggests inputs should ideally only ever be images. The discussion briefly touches on Google AI's quantum computing announcement, where they achieved a verifiable algorithm execution 13,000x faster than classical supercomputers, emphasizing the progress in hardware/computation. A major point of contention covered is the 'Definition of AGI' paper, which was criticized by Michael Saxon for containing fabricated citations; co-author Dan Hendrycks apologized, explaining the issue arose from an incorrect conversion from a Google Doc to BibTeX format, which was subsequently fixed. The video also references Karpathy's 'nanochat' project, a minimal, from-scratch ChatGPT clone trainible for under $100, demonstrating the decreasing cost of entry for training smaller, specialized models. Finally, the speaker examines the DeepSeek-OCR paper's findings on compressing long contexts and the success of the C2S-Scale 27B model in discovering a novel cancer therapy pathway, suggesting that scaling laws continue to yield breakthroughs in complex domains like biology.

### DeepSeek-OCR Performance

- Achieves 20x visual context compression with 97% OCR accuracy at <10k tokens
- Outperforms GOT-OCR2.0 and MinerU2.0 on OmniDocBench
- vLLM team working on official integration.

### Karpathy's Input Modality Debate

- Pixels are better inputs than text tokens because tokens are wasteful and carry historical baggage (Unicode ugliness, security risks)
- Suggests LLMs should only ever take images as input.

### AGI Paper Controversy

- Co-author Dan Hendrycks confirms fake citations due to conversion error from Google Doc to BibTeX; admits to fixing it after public scrutiny.

### AI Infrastructure & Cost

- Training small models like nanochat costs as low as $100 (4 hours on 8xH100)
- Larger models face quadratic scaling costs, as shown by the DeepSeek-OCR paper's 33 million pages/day training setup.

### AI in Science (C2S-Scale 27B)

- Model successfully predicted novel drug interactions, validated in lab tests showing a 50% increase in antigen presentation when combining two drugs.

### Tokenization Critique

- Tokenizers are 'ugly, separate, not end-to-end' and handle emojis poorly, resulting in weird tokens; advocates for deleting the tokenizer and processing raw pixels/images directly.

![Screenshot at 00:00: Speaker begins by showing an image related to vLLM and DeepSeek-OCR, setting the context for discussing OCR and LLM efficiency.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-00-00.png)
![Screenshot at 00:11: The vLLM project tweet detailing DeepSeek-OCR's performance metrics, including 20x compression and 97% accuracy.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-00-11.png)
![Screenshot at 00:24: The performance chart from the DeepSeek-OCR paper illustrating precision vs. average vision tokens per image on OmnidocBench.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-00-24.png)
![Screenshot at 00:53: Andrej Karpathy's tweet highlighting the philosophical debate: pixels being better inputs to LLMs than text tokens.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-00-53.png)
![Screenshot at 01:01: A Drake meme format illustrating Karpathy's preference for image inputs over text inputs for complex reasoning.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-01-01.png)
![Screenshot at 02:05: A chart showing compression performance on the Fox benchmark, illustrating precision vs. text tokens per page.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-02-05.png)
![Screenshot at 03:54: Andrej Karpathy's profile showing his background and past roles at Tesla and OpenAI, relevant to his current commentary.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-03-54.png)
![Screenshot at 04:13: A Google AI tweet announcing a quantum computing breakthrough, showing 13,000x speedup over classical supercomputers.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-04-13.png)
![Screenshot at 04:50: A Google AI blog post discussing how the Gemma model helped discover a new cancer therapy pathway.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-04-50.png)
![Screenshot at 08:56: A tweet from Michael Saxon exposing the fake citations in the 'Definition of AGI' paper, showing the cover page and conflicting reference entries for proof of fabrication.](https://ss.rapidrecap.app/screens/4D-AsJ5UhF4/00-08-56.png)
