Artificial intelligence companies are quietly purchasing thousands of printed books to build training datasets for their AI models. While the practice is legal in many cases, it has sparked growing concern among booksellers, archivists and collectors who fear that rare and out-of-print editions could disappear permanently after being digitised.
Read More | Is the AI Revolution Built on Genius or on Gimmicks?
The process is straightforward but controversial. Printed books are often dismantled by removing the spine so that each page can be scanned individually. Once digitised, the physical copies are typically discarded or sent for recycling, making it impossible to restore them. Critics argue that although the digital content survives, the original books, particularly uncommon editions, are lost forever.
Training AI Models
The issue came into the spotlight after Anthropic revealed that it had purchased millions of printed books, scanned them and used the resulting digital copies to train its Claude AI models. A US court recently ruled that scanning legally purchased books for AI training was a transformative use and therefore protected under the country’s fair use doctrine. However, the company also faced a separate legal dispute over the use of pirated digital books, which ended in a $1.5 billion settlement with authors.
Growing Demand for Printed Books
The growing demand for printed books has also created a new business opportunity. ISBNdb, which describes itself as the world’s largest book database, is now offering large-scale book acquisition services specifically for AI developers. According to the company, books published before the AI boom have become particularly valuable because they contain professionally edited, human-written content that is free from the machine-generated text now increasingly found across the internet. Orders reportedly range from 1,000 books to as many as one million copies.
Read More | Are We Confiding Too Much in ChatGPT? OpenAI’s Data Reveals a Disturbing Trend
The trend is beginning to reshape the secondhand book market. Antiquarian and used booksellers across several European countries have reported receiving unusually large orders for unrelated specialist titles. Although the buyers usually remain anonymous, some booksellers believe the purchases are linked to AI companies building new training datasets.
The practice has divided the bookselling community. For some dealers, large orders help clear slow-moving inventory and provide a valuable source of income. Others worry that books with historical, cultural or research significance could be destroyed before their rarity or importance is recognised. As AI companies continue searching for high-quality training data, concerns are growing that preserving knowledge in digital form may come at the cost of losing irreplaceable physical editions.
Indrani Priyadarshini is a journalist and editorial professional specialising in technology, artificial intelligence, smart cities, green energy, and digital transformation. With over four years of experience in tech journalism and digital media, she is known for turning complex industry developments into clear, engaging, and insightful stories. Her expertise spans reporting, editorial strategy, digital publishing workflows, and in-depth coverage of emerging technologies shaping the future.
