top of page

AI Is Devouring the World’s Books: Why Rare Volumes Are Being Bought, Scanned and Destroyed

Writer: Sue Blackwell
Sue Blackwell
Aug 23
6 min read

The strangest new customer in the rare-book market may be artificial intelligence


For generations, rare books have been bought by collectors, libraries, scholars and readers who wanted to preserve what was printed on their pages. In 2026, a very different kind of buyer has entered the market: artificial-intelligence companies hungry for high-quality human writing.


A series of investigations and newly public court records has revealed an unsettling industrial pipeline. Tech companies are purchasing physical books in enormous quantities, cutting off their bindings so the pages can move quickly through scanners, digitizing the text for AI training and then disposing of the physical copies. The practice has now moved from a niche publishing controversy into a broader debate over competition, copyright, cultural preservation and who gets to own the raw material of human knowledge.



An AirTag followed the books to Las Vegas


The most vivid evidence arrived this month. On August 17, 404 Media reported that it worked with a bookseller who placed an Apple AirTag inside a rare book included in a bulk order of roughly 1,000 titles. The tracker ultimately led investigators to an Amazon facility in Las Vegas. Workers there told the publication that books were received in large shipments, their bindings were removed and their pages were scanned.


Amazon confirmed to 404 Media that it purchases books through commercial channels to improve products and services used by customers. TechCrunch and Ars Technica subsequently reported on the investigation, highlighting the extraordinary irony: Amazon, a company that began as an online bookstore, is now using physical books as input material for the artificial-intelligence era.



Why AI companies suddenly want old books


The appeal is not difficult to understand. The open internet contains an enormous amount of text, but its quality varies wildly. It is also increasingly contaminated by machine-generated material. Books published before the generative-AI boom offer something especially valuable: long-form prose that was written, edited and published by humans before synthetic text became widespread.


For a company trying to improve a language model, that makes older books unusually attractive training material. They contain specialized knowledge, forgotten vocabulary, distinctive prose styles and subject matter that may never have been digitized. Obscure and out-of-print works can be particularly valuable precisely because their contents are difficult to find elsewhere online.


That demand appears to be changing the economics of the used- and rare-book trade. The Wall Street Journal reported this weekend that dealers have encountered unusually large orders for obscure titles, sometimes from buyers whose ultimate identity or purpose was unclear. For sellers, the boom can mean welcome revenue. For preservationists, it raises a much harder question: what happens when the value of a book to a machine becomes greater than its value as a surviving physical object?



Anthropic showed how large the operation can become


Amazon is not the first AI company linked to destructive book scanning. Court filings made public this year exposed an Anthropic initiative known internally as Project Panama. The Washington Post reported in January that Anthropic spent tens of millions of dollars acquiring and destructively scanning millions of physical books as it assembled training material for its Claude models.


The records offered a rare look inside the data race behind modern AI. Rather than simply downloading text from the web, companies have been building enormous private libraries of digitized material. Destructive scanning is attractive because removing a book’s spine allows loose pages to pass rapidly through industrial scanners. It is efficient at scale, but the original volume does not survive intact.


The distinction between an ordinary mass-market paperback and a genuinely scarce historical volume matters. Not every book purchased for scanning is culturally irreplaceable, and reporting does not establish that every AI training operation is destroying unique copies. But booksellers and preservation advocates are increasingly worried about what happens when bulk purchasing systems do not adequately distinguish between plentiful used books and editions whose physical survival has historical value.



Now Washington is being asked to intervene


The controversy took another turn on August 21. More than a dozen public-interest and consumer groups urged the Federal Trade Commission to investigate AI companies’ acquisition, scanning and destruction of books. Axios first reported the letter, and CBS News separately covered the request.


The groups are framing the issue not only as a copyright or preservation dispute but as a competition problem. Their concern is that a powerful company can purchase physical source material, digitize it into a private training corpus and then eliminate the physical copy, leaving knowledge that was once available in the open market accessible primarily through proprietary systems.


Whether regulators will accept that argument remains uncertain. Purchasing a lawful copy of a book is not the same thing as stealing one, and the legal questions surrounding AI training remain complicated. But the FTC request signals that the debate is expanding beyond whether AI companies are allowed to train on books. The emerging question is whether the methods used to acquire and consolidate that knowledge can themselves reshape markets.



A preservation problem hiding inside a technology story


Books are unusual objects because they exist in two forms at once. They are containers of information, but they are also artifacts. A first edition, an annotated copy, an unusual binding or a small-run regional history can carry meaning that is not captured by extracting its words into a text file. Marginalia, paper, typography, ownership marks and physical construction can all matter to historians.


That is why the AI scanning controversy has touched a nerve beyond the publishing industry. Digitization can preserve information and dramatically expand access. Libraries have spent decades scanning fragile material for exactly that reason. The difference is that preservation projects generally treat the physical object as something worth keeping. Destructive scanning reverses that priority: the digital text is the product, and the book can become disposable once its information has been captured.


There is a practical counterargument. The world contains hundreds of millions of duplicate books, many of which have little collectible value and may eventually be pulped anyway. Destructive scanning can be faster and cheaper than photographing bound pages individually. The difficult policy problem is not whether any book may ever be cut apart. It is how a market operating at industrial scale identifies what should not be lost.


The irony at the center of the AI race


The deepest irony is that the artificial-intelligence industry is turning back to physical books because books contain something the internet increasingly cannot guarantee: sustained human thought. The same technology capable of generating limitless text is creating new demand for writing produced before limitless synthetic text existed.


That makes old books a kind of clean data reserve. Their scarcity, age and distance from the AI-generated web can increase their usefulness to companies building the next generation of models. In economic terms, material that looked obsolete in a digital world has acquired a new strategic value.


For writers and readers, the story is therefore larger than destroyed bindings. It is about a strange reversal in the history of information. For decades, technology promised to liberate knowledge from paper. Now some of the most advanced technology companies on Earth are returning to warehouses, used-book dealers and forgotten shelves because the printed past contains precisely the human material their machines need.


What happens next


The immediate pressure is likely to fall on AI companies, booksellers and regulators to establish clearer standards. Sellers may become more selective about large anonymous orders. Preservation groups may push for safeguards around scarce editions. AI developers may face growing pressure to explain how books are sourced, scanned, retained and discarded.


None of this means digitization itself is the enemy. A scanned book can outlive a damaged physical copy and make inaccessible knowledge searchable around the world. The controversy is about what is lost when preservation is not part of the process—and about whether private AI training libraries should become the final destination for knowledge that once circulated publicly.


The race to build smarter machines has produced plenty of futuristic imagery: giant data centers, advanced chips and robots. One of its most consequential supply chains, however, looks remarkably old-fashioned. It begins with a shelf of books.


Sources & further reading








Image credit


Header photo by Daniel Brzdęk on Unsplash. The photograph is free to use under the Unsplash License. Attribution is included as a courtesy to the photographer.


Comments


© 2026 by FOURTH STREET ENTERPRISES

bottom of page