Dutch Bookseller Warns AI Firms Are Buying and Destroying Thousands of Books
TestNews Desk
Saturday, August 1, 2026
A Dutch bookseller thought a bulk order for 3,000 books was spam or phishing. It turned out to be a genuine request from an AI data company planning to scan the books and then destroy them. The revelation has renewed concerns about how copyrighted physical works are being consumed to train large language models.
A suspicious request from nowhere
When a Dutch bookseller received an unsolicited email asking for 3,000 books, his first instinct was to delete it. The message was vague, the quantities were absurd for an independent shop, and the sender claimed to represent a company that needed the books for "digital processing."
"I thought it was spam or phishing," the bookseller later told a Dutch trade publication. "No real customer orders 3,000 books without warning."
But a follow-up call changed his mind. The caller explained that the books would not be sold to readers. They would be scanned, turned into text data for artificial intelligence training, and then destroyed. The order was not a scam; it was part of a fast-growing but little-known procurement pipeline that feeds physical books into AI systems.
The bookseller ultimately refused to supply the books. But his public account, shared across the Dutch publishing world and later by international technology writers, has pulled back the curtain on a practice that many in the book industry are only now beginning to understand.
When 'spam' turns out to be real
The bookseller, who spoke on condition of anonymity to avoid retaliation from the buyer, described the process as surreal. The company wanted a broad inventory of recently published nonfiction and fiction titles, arranged by genre and topic. They offered to pay above the cover price and to cover shipping. They also said they would handle the books themselves, meaning the shop would not have to do anything except let them into the storeroom.
He suspected a phishing attack designed to harvest bank details. But after the phone call, he said, he realized it was a legitimate business proposal.
"They told me the books would be cut apart, scanned, and recycled," he said. "There would be no physical copies left for resale. They were very matter-of-fact about it."
That exchange was not an outlier. In recent years, several technology companies and their vendors have quietly purchased large numbers of physical books from bookstores, libraries, and wholesalers, often for the explicit purpose of scanning and discarding them.
The practice is sometimes called "book chopping" or "book shredding," and it has become a controversial strategy for obtaining high-quality training data for large language models.
Why AI companies want your books
AI models like OpenAI's GPT, Anthropic's Claude, and Meta's Llama rely on enormous amounts of text to learn grammar, facts, reasoning patterns, and even subtle tone. Much of that data has been scraped from the open web, including Wikipedia, Reddit, news articles, and niche forums. But web text has limitations: it is repetitive, uneven, and often shallow.
Books are different. They provide long, coherent, well-edited prose that is far better at teaching a model to handle complex narratives and arguments. According to researchers, books represent some of the highest-quality public text still available in large quantities.
For years, companies used pirated collections like Books3, a dataset that included more than 180,000 copyrighted titles. That led to lawsuits from authors and publishers, including a high-profile case brought by John Grisham, George R. R. Martin, and others against OpenAI owner Oath and its partner Microsoft.
In response, several AI companies have tried to build their own training corpora through licensing agreements with publishers and academic database providers. But licensing is expensive, and many major publishing houses remain reluctant to allow their backlists to be used for model training.
So some technology firms have turned to physical copies. By buying a book off a shelf, scanning it in-house, and then destroying the original, they can argue they are using a lawful source for digitization. That technical legality, however, has not resolved deeper copyright and ethical questions.
A hidden industrial process
The requests that reach bookstores are often processed by third-party data suppliers, not by the AI companies themselves. These vendors are hired to acquire physical books, strip off the covers, cut off the bindings, and feed the pages through high-speed scanners. After the scanning process, the paper is usually shredded or pulped.
In some cases, the text is run through optical character recognition software to make it machine-readable. Then the resulting digital file is handed over to an AI developer, sometimes without any confirmation that the original text is under copyright.
One Dutch publisher who asked not to be named said the arrangement felt like a dystopian reversal of the traditional book trade. "We have always thought of publishers as makers of books," he said. "Now there is a parallel industry where books are regarded as disposable raw materials. They are mined, not read."
That sense of unease is shared by many booksellers. A printed book, especially a hardcover or a limited edition, has cultural and material value. Destroying it for a digital model that may never credit the author feels, to many, like a kind of vandalism.
Copyright and the fair use fight
The legal status of using copyrighted books to train AI is one of the most volatile questions in tech law. AI companies have generally argued that their use of copyrighted text is "fair use" because the model does not reproduce the original text but rather learns statistical patterns from it. Authors and publishers disagree, arguing that creating a commercial foundation model from unlicensed books is no different from piracy.
Several court cases are now winding their way through the U.S. legal system, and earlier rulings have sent conflicting messages. In one prominent case involving the Google Books scanning project, a court found that digitizing books for search and snippet display was fair use. But that decision was limited to a noncommercial, publicly beneficial purpose. AI training is commercial, and the output can be direct replacement for the book itself.
Copyright scholars say the physical destruction of books adds another layer to the debate.
"The destruction of a physical copy does not extinguish the copyright in the underlying work," noted one European copyright lawyer who follows AI litigation. "Scanning the entire text and using it to train a model is still a reproduction of the work. Whether it is allowed depends on the jurisdiction and the specific use."
In the European Union, the Copyright in the Digital Single Market directive imposes additional restrictions on text and data mining. Research institutions may mine copyrighted works for noncommercial purposes, but commercial AI developers require explicit permissions unless they already have lawful access to the work. A purchased physical book is not necessarily">Legal access" in the digital sense, which may complicate the practice for companies based in Europe.
The Dutch bookseller said he was never asked about copyright. "They wanted the physical objects," he said. "They knew the copyrights belonged to someone else. They just didn't seem to think it was their problem."
How many books are being destroyed?
No one knows precisely how many physical books have been consumed by AI training. The demand is not limited to one company or one country. Wholesalers have reported orders in the tens of thousands, and some logistics firms have built dedicated facilities where pallets of books are unloaded, scanned, and shredded within hours.
One European book distributor said nearly half a million books had been purchased by AI data firms in the past eighteen months. Another source estimated the true number could be in the millions, since many orders pass through intermediaries that do not disclose their end client.
Independent bookshops have become an unlikely front line. Because major wholesalers often have contracts with publishers and may be cautious about selling to data firms, AI vendors have started approaching smaller stores directly. Their orders are large enough to look attractive to shop owners struggling with rising rents and declining foot traffic.
The Dutch bookseller's refusal seems to be relatively rare. Trade sources say many shops across the Netherlands and Germany have accepted such orders, seeing them as a lifeline for businesses that can no longer compete with online retail.
"It is hard to say no when someone offers to buy your entire stock," said a bookstore manager in Amsterdam. "But then you start to wonder if you are helping to erase the culture you are trying to sell."
A jarring economic incentive
The economics of book destruction are straightforward. AI companies are desperate for text, and books are abundant. A used book that sells for a few euros in a charity shop can, after scanning, become part of a training set that helps build a product worth billions.
For owners of unsold or old books, an AI buyer can be a godsend. They avoid shipping costs, reduce inventory, and get paid for items that were otherwise destined for the pulper. Some literary agents have expressed sympathy for that view, pointing out that remaindered and unsold books are destroyed by publishers every year even without AI involvement.
But for authors, the practice feels like an end run around the copyright system. If an AI company purchases a single copy of a book, it can only scan that copy once. But the resulting digital text can be used to train an infinite number of models, each of which might generate text that competes with the original author's livelihood.
"The value of one physical book is minuscule," said a publishing law expert. "The value of the digitized version is enormous. The system has not caught up to that discrepancy."
What happens next
Book industry groups in the Netherlands, Germany, and the United Kingdom are now organizing to warn members about AI data suppliers. Some trade bodies have drafted guidance on how to identify suspect requests and how to negotiate terms that would prevent book destruction.
One proposal, floated by a Dutch book distributor, is to require AI companies to deposit a digital copy of every scanned book into a central archive accessible to rights holders. That way, authors could at least see how their works are being used.
Another idea is to replace physical destruction with true licensing agreements, where publishers authorize the digitization of their titles for a fee and the physical books remain intact. Microsoft, OpenAI, and Anthropic have already signed licensing deals with several news publishers, and similar agreements could be extended to book publishers.
But progress has been slow. Most book publishers lack the collective bargaining power of large media groups, and individual authors have no way to negotiate with AI companies on equal footing.
The Dutch bookseller, meanwhile, says he will be more cautious about bulk orders in the future. He has started asking every unusual client about their intended use, and he keeps a list of suspicious companies to share with fellow booksellers.
"If a real reader wants three books, I will sell them," he said. "If someone wants 3,000 books and says they will feed them into a machine, I do not want that money. There is no price for that kind of destruction."
His experience may become a common story as the AI industry continues to devour the printed word. The question is whether legislators, publishers, and the public will act before the bookshelves are empty.
Comments (0)
No comments yet. Be the first to share your thoughts.
Loading stories...