“…In Bartz v. Anthropic, Judge William Alsup ruled that purchasing a book, scanning it, destroying the physical copy, and keeping only the digital version for AI training qualified as fair use under U.S. copyright law…”

From: AI firms are buying up old books because they are the last slop-free data left

The pitch from data broker ISBNdb is blunt: the world’s best AI training data is sitting on a shelf. The catch is that getting at it means slicing the spine off millions of books, then never admitting which lab paid for the job.

AI has a pollution problem of its own making. So much of the web is now machine-written that models risk feeding on their own exhaust. One company thinks the antidote is sitting on a shelf: old, printed books, published before the chatbots arrived.

***

From: AI companies are still buying up old books by the pallet – then shredding them

Researchers call it “model collapse.” Feed a language model its own synthetic output and quality deteriorates fast. Meanwhile, some writers are deliberately poisoning web content to corrupt scraping pipelines. Old printed books sidestep all of this. Edited by humans, published through traditional workflows, and existing entirely offline, they remain untouched by bots or sabotage strategies. BookData.ai, a firm marketing book-sourced datasets, describes them as “the highest-density source of structured, coherent human thought,” according to its website. For readers curious how these datasets ultimately surface in consumer tools, a look at AI-Powered Websites illustrates the downstream results.

The industrial mechanics are precise:

Standard pallets hold 800–1,200 books; buyers scale from pilot orders to 10,000+ volumes per batch, with destructive scanners processing 80–120 pages per minute after spines are cut and originals pulped afterward
Buyers deploy AutoBuy flags on ISBN lists through platforms like Alibris and Biblio - picture a Spotify playlist auto-adding tracks, except the vinyl gets shredded at the end
Anthropic‘s “Project Panama” spent tens of millions on this exact pipeline, using contractor Datamation to scan books for Claude’s training data, with booksellers identifying these buyers by abnormal volume, subject-agnostic orders, and total indifference to pricing, according to court documents reported by the Washington Post

One Court Ruling Changed Everything

A federal judge called the buy-scan-destroy pipeline “clearly transformative” – and other labs immediately took notes.

In Bartz v. Anthropic, Judge William Alsup ruled that purchasing a book, scanning it, destroying the physical copy, and keeping only the digital version for AI training qualified as fair use under U.S. copyright law. What the ruling means in practice is straightforward: buy a book, scan it, destroy it, keep the data. Because the digital copy “replaced” the physical one and was never redistributed as a book substitute, the use was deemed transformative. Anthropic separately settled for a reported $1.5 billion covering roughly 500,000 works, according to Dataconomy – preserving every trained model built on the data.