AI learns from antiquarian books, then destroys them
Translated from Slovak, summarized and contextualized by DistantNews.
At a glance
- AI companies are increasingly sourcing training data from physical books due to internet content limitations.
- Companies like Anthropic are buying, scanning, and destroying millions of books for AI development.
- This practice has led to lawsuits from authors and raises concerns about the preservation of physical literature.
Artificial intelligence companies are turning to physical books, often sourced from secondhand bookstores, to train their algorithms, a practice that involves the destruction of the original texts. As internet content becomes more restricted by paywalls and copyright laws, and increasingly filled with AI-generated material, companies are seeking alternative, high-quality text sources.
Companies such as Anthropic, known for its Claude chatbot, have reportedly spent tens of millions of dollars acquiring and scanning millions of books. Internal documents, like those from Anthropic's "Project Panama," reveal a strategy to "destructively scan all the books in the world." This process involves cutting off the books' bindings, high-speed scanning of pages, and then discarding the paper, not for archival purposes but as training material for AI.
Project Panama is our effort to destructively scan all the books in the world.
This method has not gone unnoticed, leading to legal challenges from authors who argue their copyrighted works are being used without permission. The article highlights the irony that while physical books are essential for training AI, the process itself leads to their destruction. The sourcing of these books from antiquarian dealers has also raised questions, with some booksellers now aware of the buyers' true intentions.
The practice underscores a growing tension between the advancement of AI technology and the preservation of physical cultural heritage. While AI developers seek vast datasets to improve their models, the methods employed are leading to the irreversible loss of unique physical texts, prompting debates about intellectual property, digital archiving, and the future of literature.
This document is visible to all Anthropic employees, but you should avoid talking about it in public, and the fact that we are working on it should not be shared with anyone outside Anthropic.
Originally published by SME in Slovak. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.