DistantNews
Support us
๐Ÿ‡ฎ๐Ÿ‡ฑ Israel /Technology

AI burns books? Tech giants are buying, shredding books to train smarter chatbots

From Jerusalem Post · () English

Summarized and contextualized by DistantNews.

At a glance

News Sources not specified Context piece
  • AI companies are purchasing millions of rare books to scan and shred, using the content to train smarter AI models and prevent "slop" or "model collapse."
  • Companies like ISBNdb are facilitating these large-scale book acquisitions, which raise concerns among rare book sellers due to the unusual demand.
  • This practice has led to legal challenges, including a multi-million dollar class action lawsuit by authors against Anthropic for allegedly using pirated book content without permission.

Artificial intelligence companies are engaging in a controversial practice of buying and destroying millions of rare books. The primary goal is to scan the content for training new AI models, aiming to improve their writing quality and prevent the degradation of AI performance known as "model collapse."

AI companies are buying and then destroying millions of rare books to prevent AI slop content that consumers are vocally against.

โ€” 404 MediaReporting on the core practice and its stated aim of preventing poor AI output.

Tech giants are reportedly paying companies and contractors to acquire vast quantities of books, which are then processed by high-speed machines that remove their spines before shredding the originals. Firms like ISBNdb, which claims to host the world's largest book database, argue that pre-2022 books are ideal for AI training as they are guaranteed to be free of AI-generated text. This practice has alarmed rare book sellers, who note that collectors typically buy only one or two books at a time, making the AI companies' demand for hundreds or thousands of copies highly unusual.

The world's best AI training data is sitting on a shelf. Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.

โ€” ISBNdbExplaining the rationale for using older books as training data.

Executives believe that access to extensive troves of books is essential for providing new data to AI models and mitigating the risk of "model collapse," a phenomenon where AI performance deteriorates over time due to training on lower-quality, AI-generated outputs. The Washington Post reported in January on Anthropic's efforts to acquire millions of books for this purpose.

In the rare book trade, itโ€™s very seldom that people want to buy more than one book. So if somebody comes and says, โ€˜I want a couple of hundred of your books,โ€™ itโ€™s very strange.

โ€” Pieter de VriesA rare book seller expressing surprise at the scale of AI companies' book purchases.

This trend has drawn significant attention and sparked legal action. Authors have filed a multi-million dollar class action lawsuit against Anthropic, a company backed by Amazon and Alphabet. The lawsuit alleges that Anthropic used pirated versions of their books without authorization to train its chatbot, Claude, raising serious questions about copyright and fair use in the age of AI development.

The AI companiesโ€™ attempt to hoover up books, art, news articles, and other forms of media has gotten attention in several instances since 2024.

Summarizing the broader trend of AI companies acquiring various media for training.
DistantNews Editorial

Originally published by Jerusalem Post. Summarized and contextualized by our editorial team with added local perspective. Read our editorial standards.