Millions of books died so Claude could live | The Vergecast
Quick Overview
Anthropic's Project Panama involved destructively scanning millions of books, often by slicing off the spines, to train its Claude AI models, a process revealed through newly unsealed court documents detailing initial reliance on pirated data from shadow libraries like LibGen before transitioning to bulk-purchased used books.
Key Points: Anthropic's Project Panama aimed to "destructively scan all the books in the world," involving slicing off spines for efficient digitization to train AI models powering Claude. Initial data acquisition involved piracy, with a former OpenAI executive allegedly downloading the entirety of the shadow library LibGen before repeating the action at Anthropic. Anthropic hired Tom Turvy, who oversaw the Google Books project, to manage the large-scale physical book scanning operation, bypassing slower, non-destructive methods. The acquisition of books was deemed necessary because licensing was too expensive and slow, leading Anthropic to purchase bulk quantities from used book warehouses like Better World Books. Judges in both the Anthropic and Meta cases ruled that the training of AI models on book content was fair use, though the reasoning differed significantly between the two rulings. Anthropic ultimately settled with authors for $1.5 billion over the books they acquired but did not use in commercially released models, as the judge ruled the unused data could not qualify for fair use. The reflexive backlash against AI stems partly from the 'original sin' initiated by OpenAI's cavalier approach to sourcing data, forcing competitors like Anthropic and Meta to follow similar shortcut playbooks to compete in the race to build superintelligence.
Context: The Vergecast episode features host David Pierce discussing two main topics: first, an interview with Will Arnes of The Washington Post regarding Anthropic's massive book digitization effort known as Project Panama, and second, a segment with Julia Alexander concerning Netflix's evolving strategy regarding movie theaters and studio acquisitions like Warner Brothers Discovery. The discussion on AI centers on copyright lawsuits and the ethical implications of training large language models on copyrighted material without explicit permission.