Dutch Bookseller’s 3,000-Copy Order Was Real: AI Firms Are Destroying Books

T

TestNews Desk

Saturday, August 1, 2026

A Dutch bookseller who dismissed a bulk request for 3,000 used books as “spam or phishing” has learned that the buyer was an AI data-collection firm. The company planned to scan every page and then pulp the volumes, a practice that is quietly expanding in the AI industry. The order highlights how physical books have become a raw material for training large language models.

A bulk order that looked like a scam

When an email arrived at a Dutch bookstore asking for 3,000 copies of a mix of non-fiction and literary titles, the manager assumed it was a phishing attempt. The sender was vague about how the books would be used, offered to pay in advance, and said the condition of the volumes did not matter. It was the sheer size of the order that felt wrong. A single customer had no plausible reason to need thousands of identical or near-identical copies of aging titles. The bookseller initially ignored the message, convinced that clicking on the attached spreadsheet would lead to a compromised account or a wire-transfer fraud.

But the request did not go away. A follow-up email came from a company with a professional website, a corporate registration number, and what appeared to be legitimate contracts with data-processing vendors. The bookseller decided to investigate. What he found changed his view: the company was not a library, not a school, and not a collector. It was an AI data firm that had been contracted to compile text corpora for training large language models. The 3,000 books were intended for scanning and then destruction. After the pages were digitized, the paper would be recycled.

The bookseller eventually confirmed the order with the buyer, but by then he had grown uneasy about the wider practice. He is far from alone. Across Europe and North America, booksellers, thrift stores, and remainders dealers have reported receiving bulk inquiries from companies that describe themselves as “data services” or “content digitization” firms. The buyers are often intermediaries for AI developers, and they are looking for the same thing: large volumes of text that are not already freely available on the open web.

From bookstore to training-data pipeline

Books have been used to train AI since the earliest neural language models. What is new is the scale and the physical component. Previously, many companies relied on pirated digital book collections such as the “Books3” dataset, which was assembled from a torrent site and used by researchers and later by major AI firms. That approach triggered copyright lawsuits from authors including Sarah Silverman, Paul Tremblay, and Ta-Nehisi Coates, and it pushed the industry toward alternative supply channels.

One alternative is to buy physical books in bulk, scan them in-house, and then discard the originals. The economics are straightforward: high-quality, professionally edited prose remains more valuable to an AI system than the unpredictable text of social media comments or automatically generated webpages. Books contain coherent arguments, narrative structure, and a broad vocabulary. They also reflect the cultural knowledge and stylistic conventions of decades, sometimes centuries, of writing.

A high-speed production scanner can digitize a 300-page book in a few minutes. The process typically involves cutting off the spine, feeding the loose pages through an automatic document feeder, and applying optical character recognition to produce a searchable text file. Once the scan is complete, the paper is baled and sent to a recycling plant. From the perspective of the data vendor, the physical book has served its purpose. It will never be sold again, borrowed, or read by a human.

The Dutch bookseller told local media that the buyer had initially asked for the books to be shipped directly to a warehouse, and that the covers would be removed before scanning. In a later email, the buyer said the company was “not interested in the physical object” but only in the intellectual content. The bookseller had to decide whether to participate. He eventually declined, not because the payment was doubtful, but because he did not want to be part of a pipeline that destroys books.

Why AI companies need physical books

The rise of digital books and e-readers has not reduced the demand for paper. In fact, it has created a strange new market for second-hand books that are no longer in print and cannot be licensed from publishers. Older books, especially non-fiction works from the late twentieth century, contain specialized knowledge that is underrepresented in existing training datasets. Some of those books have never been digitized legally, or were scanned so badly that their text is unusable. For AI companies, these physical copies are irreplaceable.

There is also a growing concern among AI researchers that language models have already consumed most of the high-quality text freely available on the internet. A 2024 research paper estimated that the public web is approaching a “data frontier” where new text is increasingly generated by AI itself, creating a feedback loop of lower-quality content. Physical books represent a vast, untapped reserve. The copyright clearinghouse that was supposed to regulate such use, based on voluntary licenses and public records, has not kept pace with the scale of modern AI training runs.

This has led to a legal gray area. Buying a book legally and scanning it for personal use is generally allowed in many countries under private-copy exceptions. But scanning an entire book for commercial AI training, and then destroying the original, is not clearly sanctioned by law. Courts in the United States and the European Union are still deciding whether such use qualifies as “fair use” or “text and data mining” under new directives. Meanwhile, AI companies argue that they are not copying the expressive essence of the books; they are merely extracting statistical patterns from them.

That argument does not reassure authors or publishers. “It is a tragic irony that books, which are designed to transmit knowledge across generations, are being torn apart and thrown away to be converted into a statistical model,” said a copyright law professor with experience in digital-media disputes. “The machine doesn’t read them in any human sense. It consumes them.”

The copyright and ecological questions

The destruction of physical books raises questions that go beyond intellectual property. Many booksellers view their inventory as a cultural resource, not just a commodity. Used-book stores often stock titles that have been out of print for decades, and they are the only places where researchers can find certain editions. A single bulk order can eliminate a significant portion of a niche subject area from the used-book market. If those copies are pulped, they cannot be resold to a student, a historian, or another reader.

There is also an environmental dimension. Pulping and recycling paper require energy and chemicals, even if the material is diverted from a landfill. While recycling is better than burning, the process is still a net loss of a product that was created from virgin fiber, printed, bound, and shipped. Environmentalists point out that the AI industry’s appetite for data is not purely digital; it has a physical pollution footprint that is rarely disclosed in corporate sustainability reports.

The Dutch bookseller was particularly struck by the fact that the buyer did not care about the condition of the books. “They said the covers could be dirty and the pages could be yellowed. The only thing that mattered was that every page was present,” he explained to a local technology news outlet. That indifference to the physical object is what convinced him that the books would not be read or preserved. They were destined for a scanner and a shredder.

What happens next

AI companies are now seeking more direct partnerships with publishers to avoid the legal problems of covert book scanning. Several major publishers have signed licensing agreements with AI developers, allowing their catalogs to be used as training data. But those agreements cover only new books, and only books whose rights are held by the publisher. The vast majority of twentieth-century non-fiction remains out of print, or has reverted to the authors or their estates. For that corpus, physical copies are still the only reliable source.

The Dutch book trade has begun discussing a voluntary code of conduct. Some booksellers now ask every bulk buyer to sign a written statement saying the books will not be used for AI training, or that they will be donated to libraries after scanning. Others have simply stopped responding to anonymous bulk inquiries. The seller in this case said he hopes the story will serve as a warning to other independent bookshops that are struggling to stay open in a post-Amazon world. A large order can look like a lifeline, but it can also turn a bookstore into an accomplice in the destruction of the very culture it exists to preserve.

For now, the process continues quietly. There is no public registry of how many books have been bought and scanned for AI training, and no way to know how many volumes have been reduced to pulp. The Dutch bookseller’s 3,000-copy order was unusual only because someone talked about it. The same transaction is likely happening in warehouses and back offices all over the world.

Comments (0)

No comments yet. Be the first to share your thoughts.

Loading stories...