
Secondhand booksellers across Europe and Australia are reporting a strange new kind of customer: one that buys large batches of obscure, unrelated titles, pays without much concern for price or postage, and offers little explanation of where the books are going.
The pattern has prompted a disturbing theory. As AI developers compete for high-quality training material, physical books that never became reliable digital text may be passing through intermediaries, scanning operations, and private datasets. The orders themselves are real. But the final buyer, intended use, and fate of many books remain unconfirmed.
Orders that do not look like libraries
Kennys Bookshop in Galway reportedly received an order for 5,000 books spanning subjects with no obvious relationship. German booksellers have described similarly odd requests. In Australia, multiple sellers told the Guardian that a Canadian buyer placed unusually large, price-insensitive orders for niche and decades-old stock.
There are ordinary explanations for some bulk purchases. Institutions build collections, dealers arbitrage inventory, and book recyclers aggregate stock. Zoom Books, the company named by several Australian sellers, told the Guardian that it resells books intact and does not digitize them. It declined to discuss customers because of commercial confidentiality.
That distinction matters. A suspicious order is not proof of an AI-training contract, and no responsible account should turn booksellers’ concerns into a confirmed chain of custody. What the reports establish is a recurring demand pattern. What they do not establish is a single buyer or purpose behind every order.
Why physical books have become valuable data
The open web is abundant, but abundance is not the same as quality. It contains duplicated pages, synthetic text, shallow summaries, broken formatting, and material whose licensing is fiercely contested. Older specialist books can offer something different: edited, structured knowledge that may be absent from common digital corpora.
That makes the secondhand market a possible data-acquisition layer. ISBN databases can identify missing titles. Marketplaces can locate copies across thousands of independent sellers. Intermediaries can consolidate shipments. Industrial scanners can turn paper into machine-readable text. The resulting pipeline resembles procurement more than web crawling—and is much harder for authors, publishers, and booksellers to observe.
There is documented precedent for destructive scanning. Court records in the Anthropic authors’ litigation described the company buying millions of print books, removing their bindings, scanning them, and discarding the physical copies. In that case, a US federal judge treated the scanning of lawfully purchased books for an internal research library as fair use, while separating that question from Anthropic’s use of pirated copies.
A book is also an artifact
For a mass-market paperback with thousands of surviving copies, destructive scanning may look like a procurement and copyright question. For a scarce local history, annotated volume, out-of-print technical manual, or unusual edition, it can also become a preservation question.
The text is not always the whole object. Marginal notes, bindings, inserts, ownership marks, printing details, and the relationship between editions can carry historical meaning. A dataset preserves extracted symbols; it does not necessarily preserve the artifact or even record that a particular copy once existed.
Not every bookseller shares the alarm. Istanbul antiquarian bookseller Nedret İşli told Anadolu that he found little evidence of AI companies broadly emptying European antiquarian shops and questioned the commercial logic of destroying genuinely rare material. That skepticism is useful: it separates a plausible emerging supply chain from the stronger, still-unproven claim that the rare-book trade is being systematically stripped.
The data supply chain needs provenance too
AI companies increasingly talk about the provenance of model outputs. The same discipline should apply to model inputs. A credible acquisition record would identify who sourced a work, the rights basis for using it, whether the copy survived digitization, and whether culturally significant material received preservation review.
Marketplaces and intermediaries could support that record without exposing every commercial detail. Bulk buyers could disclose acquisition categories, flag destructive scanning, exclude rare or unique copies, and provide sellers with a meaningful choice. Model developers could publish aggregate sourcing and retention policies rather than leaving the physical layer invisible.
The striking development is not simply that AI systems want more books. It is that the search for scarce human knowledge may be producing an opaque logistics industry around objects that were previously overlooked. The web made data collection feel intangible. The secondhand bookshelf makes its material cost impossible to ignore.
Relevant links
- The Guardian: Australian booksellers raise alarm over rare titles and the AI supply chain
- Tom’s Hardware: European bookshops report suspicious bulk orders
- Anadolu: A Turkish antiquarian bookseller questions how widespread the trend is
- SunMarc: Anthropic’s AI Copyright Settlement Redraws the Cost of Training Data
- SunMarc: Twitch Turns Creator Streams Into Default AI Training Data