# Millions of books died so Claude could live | The Vergecast

Source: https://www.youtube.com/watch?v=laIUf_1peuw
Recap page: https://rapidrecap.app/video/laIUf_1peuw
Generated: 2026-02-03T13:32:59.779+00:00

---
## Quick Overview

Anthropic's Project Panama involved destructively scanning millions of books, often by slicing off the spines, to train its Claude AI models, a process revealed through newly unsealed court documents detailing initial reliance on pirated data from shadow libraries like LibGen before transitioning to bulk-purchased used books.

**Key Points:**
- Anthropic's Project Panama aimed to "destructively scan all the books in the world," involving slicing off spines for efficient digitization to train AI models powering Claude.
- Initial data acquisition involved piracy, with a former OpenAI executive allegedly downloading the entirety of the shadow library LibGen before repeating the action at Anthropic.
- Anthropic hired Tom Turvy, who oversaw the Google Books project, to manage the large-scale physical book scanning operation, bypassing slower, non-destructive methods.
- The acquisition of books was deemed necessary because licensing was too expensive and slow, leading Anthropic to purchase bulk quantities from used book warehouses like Better World Books.
- Judges in both the Anthropic and Meta cases ruled that the training of AI models on book content was fair use, though the reasoning differed significantly between the two rulings.
- Anthropic ultimately settled with authors for $1.5 billion over the books they acquired but did not use in commercially released models, as the judge ruled the unused data could not qualify for fair use.
- The reflexive backlash against AI stems partly from the 'original sin' initiated by OpenAI's cavalier approach to sourcing data, forcing competitors like Anthropic and Meta to follow similar shortcut playbooks to compete in the race to build superintelligence.

**Context:** The Vergecast episode features host David Pierce discussing two main topics: first, an interview with Will Arnes of The Washington Post regarding Anthropic's massive book digitization effort known as Project Panama, and second, a segment with Julia Alexander concerning Netflix's evolving strategy regarding movie theaters and studio acquisitions like Warner Brothers Discovery. The discussion on AI centers on copyright lawsuits and the ethical implications of training large language models on copyrighted material without explicit permission.

## Detailed Analysis

The primary focus is on Anthropic's Project Panama, revealed through court documents, which sought to digitize a staggering amount of books to improve Claude's model quality, as books were deemed high-quality, vetted content essential for catching up to larger rivals like OpenAI. The process was brutal, involving destructive scanning where book spines were sliced off, with recycling trucks handling the remains. Initially, Anthropic, like others, allegedly sourced data via piracy, downloading shadow libraries such as LibGen, a practice allegedly started by a co-founder who previously worked at OpenAI. When Anthropic sought physical books, they hired an expert from the Google Books project, bypassing licensing negotiations due to cost and speed constraints, instead buying millions of used books in bulk from warehouse suppliers. Legally, judges have found the *training* on the book content to be transformative fair use in separate cases involving Anthropic and Meta, but the acquisition method remains problematic; Anthropic settled for $1.5 billion because the judge ruled that books acquired but never used in commercial models could not be protected under a fair use defense. Host David Pierce theorizes that the intense backlash against AI stems from OpenAI's 'original sin'—starting the commercial race with a cavalier approach to sourcing data, compelling everyone else to take similar shortcuts under the existential pressure to win the AI race first, even if it meant breaking eggs or ignoring ethics.

### Phone Shopping Interlude

- David Pierce is experimenting with switching from iOS to Android due to poor iPhone 16 battery life and seeks listener advice on new phone choices, including foldables.

### Project Panama Details

- Anthropic initiated Project Panama to "destructively scan all the books in the world" by cutting off spines for efficient scanning, using these texts as high-quality data to advance Claude's capabilities.

### Data Acquisition Methods

- Initial data sources included pirated shadow libraries like LibGen; later, Anthropic bought hundreds of thousands of used books in bulk from large warehouses for scanning, avoiding slower licensing efforts.

### Legal Precedents on Training

- Judges ruled that using books to train AI models constitutes 'exceedingly transformative' fair use, although the judge in the Meta case criticized the reasoning in the Anthropic case while still finding the use fair due to lack of evidence regarding market harm.

### The Acquisition Controversy

- Anthropic settled for $1.5 billion because the judge ruled that data acquired but never used in commercially released models could not be protected by the fair use defense, focusing liability on the acquisition rather than the training itself.

### The AI Backlash Theory

- The host posits that intense anti-AI sentiment originates from OpenAI's initial 'cavalier' approach to sourcing data, creating a moral paradox where companies feel compelled to take shortcuts because 'losing the race' is seen as the ultimate failure.

### Netflix and Theaters Discussion

- The show pivots to Julia Alexander discussing Netflix's potential acquisition of Warner Brothers, noting Netflix's shifting stance on theatrical releases—now claiming commitment to theaters to pass regulatory hurdles, despite previously arguing against them due to lost pay-one/pay-two window profits.

