AI Startup Halts Printed Book Collection After 404 Media Report
TestNews Desk
Sunday, August 2, 2026
A company that paid people to mail in printed books for AI training has ended the program after 404 Media investigated and exposed its copyright risks. The startup claimed owning physical copies gave it the right to digitize them, but legal experts disagreed. Following the report, the company shut down its collection effort and said it would pursue licensing agreements instead.
Program Shut Down After Probe
The company, a small AI startup that had been operating quietly for several months, abruptly ended its experiment in collecting printed books from the public. The move came within hours of an investigative report by 404 Media, a technology publication known for its coverage of AI and data rights. The report found that the startup was soliciting physical copies of copyrighted books, scanning them, and adding the text to a proprietary training dataset used for large language models. A message posted on the startup's website after the report was published read: "We have decided to pivot away from the physical books program. We apologize to the community and will take time to reflect." The founder, who spoke to 404 Media on condition that his name be withheld at the time, later issued a statement saying the program was "never intended to disregard the rights of creators."
How the Book-By-Mail Program Worked
The program, which launched earlier this year, invited anyone with a shelf full of old paperbacks to get paid for mailing them in. The sign-up process was deceptively simple: a visitor would enter the title and author of a book they owned, and if it appeared on the company's wish list, they were sent a prepaid shipping box. The wish list was not public, but excerpts shared by 404 Media showed it included works by well-known novelists, nonfiction authors, and academic publishers. Once a package arrived at the startup's offices, employees would physically scan the book, run the images through OCR software, and strip out any metadata such as the owner's initials or handwritten notes. The digital file was then dropped into a training corpus that the company claimed contained "many millions of words." For their trouble, contributors received five dollars per book, with the promise of a bonus for rare first editions.
What the Company Said
The startup's public messaging emphasized that it was not scraping the internet and that it respected the "spirit of copyright." On its website, it argued that owning a physical book is equivalent to owning a "small mechanical reproduction right," a claim that has no basis in copyright law. The company also said that it would "donate the paperbacks to local libraries" after scanning, and that it would delete the original scans once the model was trained. In an email exchange with 404 Media, the founder wrote: "We believe our approach is legal, because we are not selling the books and we are not making them publicly available. The only use is internal research." He did not respond to a question about whether the company had obtained permission from a single author or publisher, and he declined to say why the wish list included books that were clearly under copyright and available for sale.
The Copyright Problem, Explained
To copyright lawyers, the company's defense was nonsense. "The moment you digitize a book, you are making a copy, and only the copyright holder is allowed to do that," said one attorney who specializes in intellectual property and was consulted by 404 Media. The attorney noted that a physical book is a "manifestation of a copyrightable work," but the right to reproduce the work does not follow the physical object. "It is the same reason that buying a vinyl record does not let you sell the songs on it," she added. What makes this case notable is that the startup was not using the books for anything remotely resembling a traditional purpose, like preservation or access. Training an AI model that may eventually generate commercial products is a derivative use, and the doctrine of fair use does not automatically cover it, even if the book is out of print. Courts have been choosy about that question, but no judge has yet ruled that mass conversion of books into AI training data is fair.
The 404 Media Investigation
The investigation by 404 Media peeled back the layers of the startup's operation. Reporters identified the company's domain registration, which was registered with anonymity protection, but mailing address on a packing slip sent to a contributor gave them a forwarder in Delaware. Through public records and a tip from a former intern, they linked that address to a startup with three employees. 404 Media also discovered that the company had reached out to "data donation" projects on other platforms—the details of which are not essential here—before deciding to try a direct-to-consumer book collection. The publication contacted the company and asked for clarification on a number of issues: who owned the books, whether any publisher had been informed, and what would happen to the scanned data. The company's responses were evasive, leading 404 Media to include the exchange in its report. A few hours after the article went online, the submission form was gone.
Why AI Companies Want Paper Books
There's a reason a startup would go through the trouble of asking people to send paper. Most of the easily accessible text on the open internet has already been ingested by the leading AI labs. Small startups are left with corporate blogs and Wikipedia, neither of which provides the narrative complexity or stylistic variety of a well-edited book. Furthermore, many books that are most valuable for AI training—some novels, niche textbooks, and regional histories—are simply not available in machine-readable form. Libraries often restrict the digitization of their collections due to copyright, and digital databases like Google Books are not openly downloadable. Physical books, by contrast, can be bought at garage sales and library book sales for pennies. The startup's mail-in program was a way to get these works into a computer, and for the cost of a $5 payment plus shipping, it could build a dataset that would otherwise require licensing negotiations with hundreds of rights holders.
Reaction From Authors and Publishers
The shutdown drew quick praise from writers' unions and independent publishers. "This is exactly the kind of behavior that forces authors to choose between their creative work and the profit-driven priorities of AI companies," said the communications director of a prominent authors' guild, in a statement emailed to 404 Media. Another publishing industry analyst said the episode would make AI firms think twice before trying to exploit a technicality. "There is no aggregate right to scan a century of literature just because you have a scanner and a spreadsheet," the analyst said. The incident also increases the pressure on courts and regulators to define what is fair use in AI training. With Congress and the Copyright Office holding hearings on the topic, a case like this provides a concrete illustration of the problems that can emerge when a young industry races to claim the world's knowledge without written permission.
What It Means and What Comes Next
The abrupt end of the program leaves several open questions. For the startup, the challenge now is to survive. It said it will explore partnerships with publishers and "independent authors who want to license their backlists." But the reputational damage may make those partnerships difficult; at least two publishers contacted by 404 Media have already said they would not work with the company. For the wider AI sector, this story is a cautionary tale. It underscores the growing tension between the need for massive training pages and the rights of content creators. Experts say the next big legal battles will focus not on whether AI can use books, but on how to fairly compensate the people who wrote them. The printed-book loophole, if it ever existed, is now closed. Every model maker would be wise to keep its processes transparent and its licenses in order before the next investigator comes knocking.
Comments (0)
No comments yet. Be the first to share your thoughts.
Loading stories...