A reported market for pre-AI text

A July 2026 investigation by 404 Media reported that ISBNdb, a book-data business, had marketed large-scale sourcing of physical books for language-model training. Its pitch was straightforward: printed books, particularly older ones, offer edited, structured writing created before generative AI became widespread online. That makes them attractive to developers seeking material with a clear human origin rather than text that may contain synthetic or heavily recycled content.

The report also described unusual demand observed by booksellers for large volumes of ISBN-bearing titles across a wide range of subjects. However, the available reporting does not publicly identify the AI companies behind particular orders, establish that every bulk purchase came from an AI developer, or prove the total number of books acquired. Those distinctions matter. The evidence supports the existence of a service marketed to AI clients and a reported rise in unusual buying activity; it does not prove a uniform industry-wide programme by named companies.

Why physical books appeal to model developers

The web remains a huge source of text, but it is increasingly difficult to assess who produced a given page, whether it is accurate, and whether it has been duplicated across many sites. Books are not automatically reliable, unbiased or legally clear. Yet they generally arrive with useful metadata, coherent long-form structure and a record of professional editing, authorship and publication.

Older and less digitised books may be especially useful to a developer because they broaden a dataset beyond the most frequently copied material on the internet. This is a question of data diversity as much as literary quality. Training on an overly repetitive collection risks reinforcing the same language patterns and factual errors. A corpus that includes specialised manuals, historical works and niche non-fiction may help a model handle a wider variety of writing styles and subjects.

That commercial logic is different from a claim that companies are buying books primarily to deny them to rivals. A scarce physical copy can be difficult to source, but destroying it does not itself create a durable exclusive right over its ideas or text. Nor does inclusion in a training dataset make a book retrievable from a model as a complete, accurate digital edition. The competitive value lies principally in lawful access to a curated corpus and the ability to process it, not in the destruction of the paper copy.

Anthropic’s documented scanning operation

There is a confirmed precedent for destructive book scanning. In Bartz v. Anthropic, a US federal court described how Anthropic spent millions of dollars buying millions of physical books, commonly used copies. The company removed bindings, scanned pages to create searchable digital files, and discarded the source books.

In its June 23, 2025 order, the court found that Anthropic’s conversion of legally purchased print books into digital copies for its internal library was fair use on the record before it. The ruling stressed the one-to-one nature of the conversion, the destruction of the original print copy and the lack of surplus copies. It treated the digitisation and use in training as highly transformative.

The same decision drew a sharp line around books acquired from pirate libraries. The court held that Anthropic’s retention of millions of pirated copies in a central library was not justified as fair use. The case later produced a major settlement over claims connected to pirated books, rather than a blanket judicial approval for every way an AI company might acquire or use books.

That distinction is central to the current reports. Purchasing a physical copy may address one problem of provenance, but it does not provide a universal answer to copyright questions. Fair-use analysis remains fact-specific, considering the purpose of copying, the work used, the amount copied and market effects. Different datasets, security practices, model outputs and commercial uses could lead to different outcomes.

Removing a binding is a commonplace industrial scanning method: loose pages can be passed quickly through high-speed equipment, whereas bound volumes are slower and harder to digitise. It is also cheaper to dispose of the original copy than to store it. Neither point removes the cultural concern when the book is rare, out of print, locally published or poorly represented in libraries.

A scan retained privately for model development is not a public preservation programme. It may capture text, but it does not necessarily preserve the physical artefact, its typography, binding, annotations, provenance or the ability for readers and scholars to inspect it. Nor does a language model function as a public archive. It generates responses from learned statistical patterns; it is not designed to supply a faithful, complete work on demand.

The risk should not be overstated. Millions of used books are routinely removed from circulation through ordinary resale, pulping and disposal, and no evidence in the reporting identifies a specific irreplaceable title destroyed for AI training. Nevertheless, bulk purchasing creates a legitimate incentive for safeguards. Vendors and intermediaries could screen for unique or unusually scarce copies, offer them to libraries or preservation groups first, and retain archival-quality scans where copyright and access arrangements permit.

The case for clearer sourcing rules

The episode highlights a mismatch between the scale of AI data acquisition and the systems built to document literary rights and library holdings. A buyer may lawfully own a copy of a book while the rights to reproduce it, distribute a scan or build a commercial training corpus remain contested. Meanwhile, authors and publishers may have little visibility into whether their works were included.

The US Copyright Office has concluded that voluntary licensing markets for AI training should be allowed to develop before government intervention is considered. That approach places importance on workable records of provenance, rights information and permitted uses. It also creates an opportunity for publishers, collecting societies, libraries and AI developers to distinguish commercial training copies from preservation copies.

For AI companies, transparent sourcing can reduce legal and reputational risk. For the book trade, it can turn an opaque purchasing rush into a more accountable market. And for libraries and archives, it is a reminder that preserving a text is not the same as feeding it into a model. The most durable response is not simply to prohibit scanning or to accept it without conditions, but to establish clearer rules for lawful access, compensation where appropriate, disclosure and preservation of culturally significant works.

Sources