The latest AI boom is creating an unusual paradox. The more AI companies need high-quality information, the more aggressively they are buying the physical books that contain it, and in some cases destroying those books once their contents have been absorbed. In recent months, independent booksellers in the United States, Britain, Europe, and elsewhere have reported sudden surges in unexplained bulk purchases, sometimes involving hundreds of unrelated and obscure titles. Some of these books are subsequently being sent to facilities where their bindings are removed, their pages rapidly scanned, and the remains discarded or recycled.
The scale could become far larger than the cases already exposed. ISBNdb, a major book database, now facilitates purchases ranging from 1,000 to one million books per order, while booksellers have described previously unsold titles suddenly becoming highly sought after. At the same time, an Amazon facility in Las Vegas has been identified as a site where employees described cutting book spines, scanning loose pages, and disposing of the resulting paper. The controversy raises a question that extends well beyond publishing. What happens if the race to build better AI systematically removes the physical sources from which that intelligence was created?
The immediate driver is straightforward. Large language models, the systems behind tools such as ChatGPT, Claude, and Gemini, require enormous quantities of text to improve their ability to generate accurate and authoritative responses. As easily available digital material becomes increasingly saturated with AI-generated content, older physical books have acquired a new value. They offer information produced before the widespread emergence of generative AI, making them less likely to contain text generated by earlier models.
This has transformed the economics of the secondhand book market. Sellers have reported unusually large orders for books ranging from obscure academic works to foreign-language texts and old government publications. One bookseller described moving from fewer than 20 books a week to hundreds. Another received a single order equivalent to roughly a week’s normal sales. For sellers holding thousands of old or difficult-to-sell titles, the financial incentive is obvious.
However, the same process creates a structural problem. A book purchased by a reader remains available to another reader. A book purchased for destructive scanning does not. Once its spine is removed and its pages are fed through an industrial scanner, the physical copy may be reduced to loose paper for recycling.
Evidence from an Amazon employee shows what this process looks like in practice. The employee described machines cutting the bindings from books before scanners rapidly photographed the pages. The facility had at least 20 to 25 scanners, while workers saw books arriving from different countries, including Germany, Russia, Japan, and Britain. Some were new, while others appeared to have come from libraries.
This matters because the value of a book is not necessarily related to how commercially popular it is. Millions of copies of a bestselling novel may be easily replaceable, but some obscure works may have only a handful of surviving copies. A rare academic publication, an out-of-print foreign-language book, or a signed edition can therefore become disproportionately vulnerable when multiple AI companies seek the same underlying material.
The problem could also become self-reinforcing. If several AI developers require similar high-quality training material, each company has an incentive to acquire its own physical copy. A rare book that could once circulate among readers, collectors, libraries, and researchers could instead be purchased repeatedly and destroyed each time. The result would not simply be digitisation. It would be the gradual removal of physical cultural resources from circulation.
The controversy has emerged partly because existing copyright rules may allow behaviour that is commercially rational but culturally damaging.
In the United States, a 2025 ruling found that Anthropic’s use of legally purchased books to train its AI system was transformative and therefore protected as fair use. Court filings in the underlying case revealed that Anthropic had purchased millions of physical books, removed their bindings, scanned them, and destroyed the resulting copies. The company later agreed to a $1.5 billion settlement concerning separately pirated works.
The legal logic creates an important distinction. Copyright law may determine whether an AI company is entitled to use a book, but it does not necessarily determine whether society should want that book to disappear afterward.
That distinction becomes increasingly important as AI companies scale their operations. The Federal Trade Commission has been asked by more than a dozen public-interest and consumer groups to investigate what they describe as “hoard-and-destroy” practices. Their concern is not only copyright; They argue that systematically removing books from public access could reduce competition by making essential source material harder for other AI developers, researchers, and members of the public to obtain.
The potential competitive implications are significant. If one company purchases the remaining copies of a scarce title and destroys them after scanning, another company may no longer be able to acquire the same material legally. The original content may effectively become concentrated inside the digital systems of whichever firms were wealthy enough to purchase it first.
This creates a paradox at the heart of the AI data economy. A technology industry that depends on the widespread availability of knowledge could gradually make that knowledge less accessible in physical form.
Other jurisdictions illustrate that this outcome is not inevitable. Singapore’s copyright framework explicitly permits AI training using lawfully accessed copyrighted material. Legal experts argued that this provides greater certainty for developers, while Google’s president of global affairs, Kent Walker, said that “certainty is what the industry needs.” At the same time, Walker pointed to Google’s earlier book-digitisation practices, where special books could be scanned carefully without being damaged.
This suggests an alternative model. AI companies do not necessarily need to destroy books to extract their contents. Non-destructive scanning, library partnerships, licensed collections, and systems designed to preserve rare material could allow companies to acquire training data without permanently reducing the world’s physical stock of knowledge.
Yet incentives matter. Destructive scanning is faster and potentially cheaper at industrial scale. That creates a risk that preservation will lose to efficiency unless companies, regulators, libraries, and booksellers establish rules that distinguish ordinary commercial copies from culturally significant ones.
The most consequential scenario is not that every book will disappear. It is that AI’s demand could systematically target precisely those books that are already difficult to replace.
This would produce a form of cultural depletion. Bestseller novels would remain relatively safe because enormous numbers of copies exist. The greater risk would fall on books that never achieved mass circulation. Some may have been printed in only a few hundred copies, while others may represent unique editions or contain specialised knowledge unavailable elsewhere.
The scale of this problem is still uncertain. One bookseller cited in the Orange County Register estimated that destructive scanning could already involve tens of millions of books, although this remains an estimate rather than an established industry-wide figure. What is clearer is that the practice is expanding beyond the specific Anthropic case that initially brought it to public attention. Amazon has been linked to large-scale book scanning, while booksellers across multiple countries have reported suspicious bulk purchases.
If the trend continues, the effects could extend beyond rare-book collecting. Libraries, universities, historians, journalists, and researchers depend on the continued availability of obscure material precisely because it is not commercially prominent. Destroying one obscure book may appear insignificant. Destroying the last few accessible copies of thousands of obscure works is fundamentally different.
There is also a second-order risk. AI systems do not necessarily preserve the books they ingest as searchable digital libraries. Training a model does not mean that the model simply stores an accessible database of every book it has read. The original physical object can therefore disappear without being replaced by an equivalent public digital archive.
That distinction could eventually create a knowledge bottleneck. Humanity would possess increasingly powerful systems trained on an enormous historical record, while simultaneously possessing fewer physical copies of the source material needed to verify, study, or reinterpret that knowledge independently.
The publishing industry faces a related problem. Booksellers are often financially penalised for rejecting suspicious orders because marketplace ratings depend on fulfilment. A seller who refuses an AI-related purchase may simply push the buyer toward another seller. This creates a collective-action problem in which individual sellers have strong incentives to sell even when they believe the broader outcome is harmful.
The market therefore cannot be expected to solve the problem by itself. A bookseller who protects a rare copy may lose a sale, while another seller may accept the order. Unless buyers are required to distinguish rare or culturally significant books from replaceable copies, preservation becomes an individual burden rather than an industry standard.
The most plausible future is consequently not one in which AI literally destroys all books. It is one in which AI gradually changes which books survive in the physical market. Common books remain abundant because they are constantly reproduced, while obscure books are increasingly extracted from circulation because their scarcity makes them disproportionately valuable as training data.
That would amount to an unintended reversal of the traditional economics of cultural preservation. Books that readers value most may be preserved, while books that AI companies value for their information content could become increasingly difficult for humans to access.
The issue also exposes a broader tension in the AI economy. Training models require converting physical and intellectual resources into computational assets. Once the information has been extracted, the original resource can appear economically redundant. Books are therefore treated less as cultural objects and more as temporary containers for data.
The danger is that this logic undervalues the option to access the original. A book is not merely the text it contains. Its edition, annotations, illustrations, physical condition, provenance, and continued availability can all have historical and research value. Once destroyed, those characteristics cannot be reconstructed from an AI model.
The emerging question is therefore not whether AI companies should be allowed to learn from books. It is whether they should be allowed to consume the physical foundations of the knowledge economy without distinguishing between replaceable inventory and irreplaceable cultural material.
AI companies are unlikely to abandon their search for high-quality training data. Demand will probably continue to increase as developers compete to build more capable models. The critical variable is how that demand is managed. If companies continue treating every purchasable book as an interchangeable training input, the market could remove scarce works from human circulation at a scale that becomes visible only after the damage is done.
The alternative is already technically possible. Rare books can be scanned without destroying them, libraries can participate in controlled digitisation programmes, and training datasets can be assembled through licensing and preservation agreements. Google’s Walker has argued that rare books should be handled carefully, pointing to the company’s earlier approach of turning pages “one by one, very carefully” when digitising special books.
The central challenge, then, is not technological capability but institutional incentives. Without safeguards, the economics of AI may reward companies for extracting knowledge as quickly as possible, even when doing so permanently reduces the physical supply of that knowledge. If that continues, AI could create an extraordinary historical paradox. The machines may become better informed precisely as the human archive they learned from becomes poorer.
The race for AI is increasingly a race for humanity’s accumulated knowledge. The question is whether winning that race requires destroying part of the archive along the way.
Cerullo, Megan. 2026. ‘AI Companies Accused of Hoarding and Destroying Millions of Books’. CBS News, August 21. https://www.cbsnews.com/news/ftc-ai-companies-destroying-books/.
Cole, Samantha. 2026. ‘Company Offering Printed Books to Train AI Stops After 404 Media Coverage’. 404 Media, July 31. https://www.404media.co/ai-company-training-scanning-books-database-isbndb/.
Convery, Stephanie. 2026. ‘“More Than Just Objects”: Australian Booksellers Raise Alarm Over “Horrific” Destruction of Rare Titles to Feed AI’. The Guardian, The Guardian, August. https://www.theguardian.com/technology/2026/aug/02/australian-book-sellers-alarm-destruction-rare-titles-ai-supply-chain?CMP=Share_AndroidApp_Other.
Edwards, Benj. 2025. ‘Anthropic Destroyed Millions of Print Books to Build Its AI Models’. Ars Technica, June 25. https://arstechnica.com/ai/2025/06/anthropic-destroyed-millions-of-print-books-to-build-its-ai-models/?ref=404media.co.
James, Kathryn. 2026. ‘Why Is Anthropic Destroying Books?’. The Guardian, The Guardian, August 5. https://www.theguardian.com/commentisfree/2026/aug/05/anthropic-ai-destroying-books.
Korn, Melissa. 2026. ‘Rare-Book Sales Are Booming. They’re Getting Sliced up and Fed to AI.’. The Wall Street Journal, August 22. https://www.wsj.com/articles/ais-need-for-content-has-put-rare-book-dealers-in-a-bind-1ac5a053.
Loffhagen, Emma. 2026. ‘Book Publishers Sue Google for Copyright Infringement Over Gemini AI Training’. The Guardian, The Guardian, July 14. https://www.theguardian.com/books/2026/jul/14/publishers-sue-google-gemini-ai-training?ref=404media.co.
Maiberg, Emanuel. 2026. ‘Inside the Warehouse Where Amazon Scans and Destroys Books for AI Training’. 404 Media, August 26. https://www.404media.co/inside-the-warehouse-where-amazon-scans-and-destroys-books-for-ai-training/.
Milmo, Dan. 2026. ‘Secondhand Booksellers in UK and Ireland Suspect AI Firms Behind “Strange” Bulk Orders’. The Guardian, The Guardian, August 15. https://www.theguardian.com/technology/2026/aug/15/uk-ireland-booksellers-suspect-ai-companies-bulk-orders-data-acquisition.
Silman, Anna. 2026. ‘AI Has Plunged the Book Publishing Industry Into Utter Chaos’. The Wall Street Journal, August 17. https://www.wsj.com/arts-culture/books/generative-ai-book-publishing-be79a287.
Wain, Philippa. 2026. ‘Secondhand Book Sales Are Booming. Is It Because of AI?’. BBC, August 15. https://www.bbc.com/news/articles/cp3rprx2wl4o.
Comments