For generations, the life of a printed book followed a familiar path. It usually moves from a bookstore to a reader’s shelf, then perhaps to a second-hand shop or a library archive.
Today, however, some books are taking a very different route. They are bought in bulk, scanned into machine-readable text and, in some cases, destroyed, leaving behind digital files that can be used to train artificial intelligence (AI) systems.
The practice may sound unusual, but it is no longer confined to speculation.
Court records, company documents and recent investigative reporting show that a commercial market has emerged around sourcing printed books for AI training.
Among the companies operating in this space is ISBNdb, which has expanded beyond bibliographic and inventory services to help clients acquire books in large quantities.
According to a recent investigation by 404 Media, bulk purchases now range from around 1,000 books to as many as one million, with buyers’ identities shielded by strict confidentiality agreements.
Many of the books are then sent to specialist digitisation firms. In some cases, the process is destructive: the spine is cut away so loose pages can pass through high-speed scanners, a method that is significantly faster and cheaper than scanning intact books one page at a time.
Companies offering these services openly advertise destructive scanning alongside non-destructive alternatives, although there is no evidence that every bulk purchase ultimately results in the physical book being destroyed.
Books published before the rapid spread of generative AI, particularly before 2022, have reportedly become especially desirable because they predate the flood of AI-generated material on the internet and are therefore regarded as cleaner sources of human-authored text.
ISBNdb itself has acknowledged the sensitivity surrounding the practice. As quoted in 404 Media, the company noted that “the optics problem is real” and that “‘AI company destroys two million books’ is not a headline that generates sympathy.”
The legal basis for parts of this activity was strengthened by Bartz v Anthropic, a copyright case brought by authors Andrea Bartz, Charles Graeber and Kirk Wallace Johnson. They alleged that Anthropic had copied their books without permission while building a digital library and training Claude, its generative AI system.
In June last year, US District Judge William Alsup ruled that using legally acquired books to train large language models was an “exceedingly transformative” use under US copyright law.
He also held that Anthropic could lawfully convert purchased print books into searchable digital copies when the physical copies were destroyed and the files were retained internally.
The ruling did not extend to Anthropic’s use of millions of books downloaded from shadow libraries such as LibGen. Those piracy claims were later resolved through a $1.5 billion settlement, preliminarily approved in September last year and finalised in July this year.
The settlement addressed only the piracy allegations. It left untouched the court’s separate finding that purchasing printed books, scanning them and using the resulting digital copies for internal AI training can fall within the scope of fair use.
Court records also shed light on the scale of Anthropic’s efforts.
In 2024, the company hired former Google Books executive Tom Turvey to lead “Project Panama”. Documents unsealed during the litigation describe an initiative to destructively scan books on an industrial scale for internal AI training.
Vendor proposals discussed processing hundreds of thousands to millions of books, while internal records indicated that the company spent tens of millions of dollars on acquisitions and scanning operations and instructed staff to keep the project confidential.
Not every claim surrounding the practice, however, has been borne out by the available evidence.
There is no public evidence linking ISBNdb or any other intermediary to specific anonymous bulk purchases, and booksellers can only infer the involvement of AI companies from unusual buying patterns.
Similarly, claims that rare or irreplaceable books have been destroyed remain unverified, with no documented examples tying identifiable scarce works to the scanning pipeline.
Even so, the existence of the pipeline itself is no longer in dispute. Companies are buying physical books in enormous quantities, converting them into machine-readable text and, in many cases, discarding the original copies.
What remains unknown is the scale of the practice. The total number of books that have passed through this pipeline, the identities of most of the companies driving it and its long-term implications for publishing, preservation and the historical record remain largely hidden from public view.