Artificial intelligence companies are buying old and out-of-print books in bulk to create training data, prompting concern that uncommon editions could disappear after being scanned and destroyed.
The process often involves removing a book’s spine, scanning individual pages and discarding or recycling the remains. Anthropic used this method on millions of legally purchased books while developing its Claude models. A US judge found that digitising purchased copies for training was transformative and protected by fair use. A separate dispute over pirated digital books ended in a $1.5bn settlement with authors.
ISBNdb, which describes itself as the world’s largest book database, is marketing acquisition services to AI developers. It says printed works from before the AI boom are especially valuable because they contain edited human writing unaffected by machine-generated material increasingly found online. Orders can reportedly range from 1,000 books to one million.
Secondhand and antiquarian booksellers in several European countries have reported unusual orders for unrelated specialist titles. Some suspect AI laboratories are behind the purchases, but buyers are anonymous.
The practice has divided traders. Bulk sales can clear unwanted stock and provide income, but booksellers fear rare or irreplaceable works may be destroyed without assessment of their cultural value.