AI companies are buying printed books by the million, slicing off their bindings, scanning the loose pages at industrial speed, and destroying the originals — and a book-data company called ISBNdb has built a business around making sure nobody knows which lab placed the order.
The practice was detailed in a 404 Media investigation published on 21 July 2026 by reporter Emanuel Maiberg, and picked up over the following week by Futurism, Tom's Hardware and The Next Web. It moved from trade-press story to mainstream backlash on 28 July, when the topic reached the top of X's AI news trends with more than 13,000 posts.
Key Highlights
- ISBNdb, which operates what it calls the world's largest book database, now sources physical books in bulk specifically for AI training pipelines.
- Single orders range from 1,000 copies up to roughly one million titles, with the buyer's identity shielded from the sellers.
- Books printed before 2022 sell at a premium because they predate the generative-AI era and are therefore free of machine-written text.
- Scanning is destructive: hydraulic cutters remove the spine, pages are fed through high-speed scanners, and the remnants are pulped or recycled.
- A June 2025 federal ruling in Bartz v. Anthropic found that buying a physical book and destructively digitising it for internal use is fair use.
Details
ISBNdb's pitch to AI labs is blunt. "The world's best AI training data is sitting on a shelf," the company's marketing reads. "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative."
The commercial logic is model collapse. As synthetic text floods the open web, the crawlable internet becomes progressively less useful as a pre-training corpus — models trained on the output of earlier models degrade across generations. Print runs from before 2022 are, as one description of the pitch puts it, structurally guaranteed to be uncontaminated. You cannot retroactively insert AI slop into a book that was already on a shelf in 2019.
That scarcity has reshaped the second-hand book trade almost overnight. A professional bookseller told 404 Media that orders surged starting in April 2026: a good week used to mean about 20 books, and recent weeks have run into the hundreds. Rare-book dealers in the Netherlands have reported the same pattern of anonymous bulk requests. Buyers frequently ask for obscure, foreign-language and out-of-print titles rather than common paperbacks — which is precisely what alarms preservationists.
The Anonymity Product
The most-quoted line from the reporting is ISBNdb's own acknowledgement of the reputational risk. "The optics problem is real," the company's material states. "'AI company destroys two million books' is not a headline that generates sympathy."
Its answer is to sell discretion as a feature: bulk sourcing under NDA, buyer anonymity toward sellers, filtering by publication year and subject, and delivery formatted for a scanning pipeline. Critics on X seized on the framing of the work as "digital preservation," noting that preservation normally implies the original survives.
The Legal Ground
The activity accelerated because a court blessed the method. In June 2025, US District Judge William Alsup ruled in Bartz v. Anthropic that training a large language model on lawfully purchased books was "exceedingly transformative" and qualified as fair use. Crucially, he also held that destructively scanning those purchased copies was a permissible format shift — the digital file replaced the physical book rather than adding a new copy in circulation.
Court documents describe Anthropic's internal effort, referred to as Project Panama, as an attempt to obtain the world's books through legitimate bulk purchase. Anthropic hired Tom Turvey, formerly of Google Books partnerships, in early 2024 to run the acquisition programme. The same ruling drew a hard line at pirated digital books, which were not covered by the fair-use finding; Anthropic later settled the piracy claims in a deal reported at roughly $1.5 billion.
The practical effect of the ruling is a legal safe harbour that rewards destruction. Keeping the physical copy weakens the format-shift argument; destroying it strengthens it.
Impact
For publishers and authors, the ruling narrows the copyright fight considerably: the purchase-and-scan route is settled law for now, and the remaining litigation centres on pirated corpora.
For libraries, archivists and the antiquarian trade, the concern is different and harder to reverse. A shredded website can be re-uploaded and a bestseller can be reprinted, but an 18th-century volume with three surviving copies cannot be un-pulped. There is no public registry of what has entered these pipelines, and the NDAs are designed to keep it that way.
For developers and teams building on top of frontier models, the story is a concrete illustration of how expensive clean data has become. The scramble for pre-2022 print is a measurable price signal on human-written text — and a preview of what the next round of pre-training will cost.
What's Next
Expect three fronts to open. Preservation bodies and national libraries are likely to push for disclosure requirements or exclusion lists for rare and low-circulation titles. Publishers may start writing scanning restrictions into supply agreements with resellers. And the fair-use reasoning in Bartz will be tested again as similar cases work through appeal, where the format-shift argument — and the incentive to destroy that it creates — is the obvious target.
Source: 404 Media