A cultural resource, not merely a dataset

Books are unusually valuable material for generative AI. They contain sustained arguments, narrative structure, specialist knowledge, dialogue, historical perspective and distinct literary voices. Compared with short online posts, a book is a dense record of deliberate human expression. That makes large collections attractive to developers seeking to improve language models’ ability to explain, reason, imitate styles or produce long-form prose.

The concern is not simply that an AI system may reproduce a passage too closely. The deeper risk is economic and cultural. If companies can extract value from vast libraries of writing without a durable mechanism to obtain permission or share returns, the incentives that fund new books may weaken. Authors, editors, translators, independent publishers and literary estates do more than supply text: they maintain a long-term ecosystem in which new voices can develop, research can be undertaken and difficult or niche subjects can find readers.

Calling this outcome a “cultural void” is necessarily a warning rather than a proven forecast. Books will not disappear because language models exist. Yet the warning identifies a real feedback loop: cultural production requires investment; investment depends on credible routes to income; and the next generation of useful training material is created by the same people and institutions whose markets may be disrupted.

The legal debate has often been framed around a binary question: is training on copyrighted books fair use, or is it infringement? In practice, the answer is more contingent. It can depend on how a work was acquired, the purpose and transformation involved, the scale of copying, the model’s outputs and its effect on current or potential markets.

Recent US litigation has illustrated this complexity. In 2025, a federal court found in the Meta books case that the plaintiffs before it had not demonstrated sufficient market harm to defeat Meta’s fair-use defence for the model-training claim. The ruling did not establish that every use of copyrighted books to train an AI model is lawful. It was tied to the evidence and arguments in that case.

The dispute involving Anthropic drew an even sharper distinction. A court treated the use of books for training differently from the company’s acquisition and retention of millions of allegedly pirated copies. In July 2026, final approval was granted for a US$1.5 billion settlement over the piracy-related claims. The episode showed that the provenance of a dataset matters independently of the question of whether a particular training use is transformative.

The US Copyright Office has likewise avoided a universal answer. Its 2025 report on generative-AI training concluded that some uses may qualify as fair use and some may not, depending on the circumstances. It also said voluntary licensing can be feasible in at least some situations. That leaves room for courts to decide individual disputes, but it also leaves authors and AI developers operating amid substantial uncertainty.

Why licensing is about more than payment

A well-designed licensing market could make the debate less adversarial, but only if it recognises what is actually being licensed. A book used to train a general model is not equivalent to a book offered through a searchable subscription service, used to generate a synopsis, translated, narrated synthetically or made available for question-and-answer interactions. Each use has different risks, commercial effects and potential benefits.

This is why author groups have argued for explicit, title-level permission rather than treating AI rights as silently included in older publishing contracts. Publishers may own or control certain digital rights, while authors may retain others. Older agreements were usually not negotiated with large-scale model training in mind. Clear consent, transparent terms and a defined share of revenue can prevent an AI licence from becoming another opaque exploitation of a writer’s work.

Licensing also has informational value. It can create records of which works entered a corpus, on what terms and for which functions. That transparency would make it easier to audit whether a model was built from authorised material, whether a work is being used beyond the agreed scope, and whether compensation reached creators rather than being absorbed entirely by intermediaries.

There are practical objections. A model developer may need material from thousands or millions of works, and negotiating one agreement at a time would be slow and expensive. Collective licensing, in which a representative body grants permissions and distributes revenue, is one possible response. It will be credible only if participation is genuinely voluntary, distribution rules are understandable, and authors can choose not to participate. A compulsory arrangement that pays minimal sums while removing meaningful control could reproduce the original problem in a more formal form.

The danger of substituting for the market that made the books possible

The most important question is not whether an AI system literally replaces a novel, biography or textbook. It is whether it substitutes for enough reading, commissioning, educational use or rights sales to reduce the financial base for producing those works.

A chatbot that provides a brief factual explanation may direct some users towards a book. It can help readers discover titles, make catalogues more accessible and assist with translation, indexing or accessibility. Used under accountable terms, these are meaningful public benefits.

But the same technology can provide extensive summaries, simulate an author’s voice, produce low-cost derivative material and saturate digital marketplaces with automated text. The likely effects will vary by genre. Reference publishing, educational materials and highly formulaic commercial fiction may face different pressures from poetry, literary fiction or specialised scholarship. The impact will also be uneven: an established author with multiple revenue streams is not situated like a debut writer, a freelance translator or a small press operating on narrow margins.

This makes market evidence essential. Arguments about harm should not rely only on the intuitive claim that a model has “read” a book. They should examine sales patterns, licensing opportunities, search visibility, reader behaviour, output similarity and the conditions under which AI-generated substitutes enter the market. Conversely, developers claiming public benefit should demonstrate that their systems create new capabilities rather than merely taking a cheaper route to content that readers would otherwise buy or borrow.

A second feedback loop: the quality of future knowledge

There is also a technical reason to value continuing human authorship. Research on recursive training has found that indiscriminately training generative models on content produced by earlier models can cause “model collapse”: rare or less represented patterns are progressively lost, and output quality becomes distorted. The phenomenon should not be overstated. Synthetic data can be useful when it is carefully generated, filtered and combined with reliable human-produced data. It is not a proof that all AI-trained systems will deteriorate.

Still, the research reinforces a cultural point. A healthy information environment needs sources that are independent of the models drawing on them. Books created through reporting, scholarship, lived experience, editorial judgement and imaginative labour offer material that is not simply a probabilistic remix of previous machine output. They preserve minority perspectives, inconvenient detail, unusual forms and long arguments that can be economically unattractive to produce but socially valuable to retain.

If original human work becomes harder to finance, the cost is not confined to authors’ incomes. Future systems may have less access to high-quality, diverse and trustworthy material. The public may also encounter a more homogenous online culture in which automated prose is abundant but the underlying pool of firsthand knowledge and distinctive expression is thinner.

Building a sustainable bargain

The sensible policy aim is neither to freeze AI development nor to treat books as free raw material because copying is technically easy. A sustainable bargain would combine provenance standards, meaningful licensing options, contractual clarity and protection against outputs that substitute for or misrepresent original works.

Publishers and authors should be able to negotiate AI uses explicitly. AI companies should be able to obtain dependable, high-quality corpora without relying on piracy or legal ambiguity. Readers should benefit from useful tools while retaining access to a living, varied book culture.

The cultural void is therefore not an inevitable technological consequence. It is a governance failure that could emerge if the value extracted from books is detached from the conditions needed to write, edit, publish and preserve the next ones.

Sources