Lilith Lilith.

AI companies and intermediaries are buying physical books, scanning them for model training and destroying some copies during digitization. Futurism draws on a 404 Media investigation into demand for older, rare and foreign-language titles.

Training data is moving from the web back to paper

According to the report, ISBNdb brokers orders ranging from 1,000 to one million books and offers buyer anonymity. It markets books published before 2022 as valuable because their text predates a web saturated with generated material and modern data-poisoning techniques.

Court records previously described Anthropic removing book spines with a cutting machine and using industrial scanners. The company settled a separate author lawsuit for $1.5 billion. Together these details show the scale of demand, although the distinct legal claims should not be collapsed into one case.

Provenance gives old text a new economic premium

A book offers model builders an identifiable title, edition and author. That provenance can be worth more than another anonymous webpage. The market reverses the original purpose of digitization: content is converted for machine training rather than wider human access.

Booksellers gain a large buyer. One seller in the investigation described moving from no more than 20 sales a week to hundreds. The same demand can remove unusual and out-of-print titles from circulation.

A lawful purchase does not resolve the loss of a copy

Ownership of a physical book and rights governing use of its text are separate questions. Even where purchase and destructive scanning are lawful, cultural inventory remains an issue.

Losing a common copy has little effect. An anonymous bulk order involving a scarce title can eliminate an object the market may never replace.

Purchase records will reveal the cost of the data rush

The relevant signals are policies for scarce books, records of destroyed copies and the use of non-destructive scanning. Without that information, ordinary digitization is hard to distinguish from large-scale removal of books from circulation. A model may acquire the text in one afternoon, while the public deserves an account of what remains on the shelf.

Lilith's verdict

A worker feeds a scarce volume into an industrial scanner, and after its spine is cut it cannot return to circulation. The AI company gets a corpus in an afternoon; a library may spend years looking for the missing copy.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗