Amazon is buying and destroying rare books to train its AI models

A stack of rare old books on a workbench, representing Amazon's practice of scanning them for AI training data.

Amazon, the company that began as an online bookstore, is now buying rare and out-of-print books and physically destroying them to feed its AI training pipeline, according to a report from 404 Media published August 17, 2026. The investigation used a tracking device placed inside a rare book, which ultimately arrived at an Amazon facility in Las Vegas identified as VGT3, a warehouse marked with a dinosaur holding a book in its claws. Amazon confirmed to 404 Media that it “purchases books through commercial channels to improve the products and services customers use.”

The practice underscores a growing challenge for AI developers: the internet’s readily available text has largely been exhausted, and companies are turning to physical archives to find new, high-quality training data. Rare books, particularly those never digitized or widely circulated, are prized because they are guaranteed to have been written by humans, not AI. This is important because training on AI-generated text risks “model collapse,” a phenomenon where an LLM’s outputs become increasingly repetitive and degraded after ingesting too much synthetic content.

Also read: How to Check If Your ChatGPT, Claude, or Perplexity Account Has Been Hacked

How Amazon’s book destruction works

According to the 404 Media report, Amazon acquires books through various channels, including bulk purchases from dealers and auctions. Once acquired, the books are sent to facilities like VGT3, where their spines are cut off and pages are scanned. The physical books are then discarded, effectively destroying them. This process, sometimes called “destructive digitization,” is not new — Google faced criticism in the past for similar practices during its Google Books project, but the scale and purpose here are different.

Amazon’s statement to 404 Media did not address the destruction directly, only noting that the company buys books to improve its products and services. The company has not specified which AI models or services benefit from this data, but the timing aligns with the broader industry push to secure exclusive, high-quality training datasets.

Also read: Groq Raises $350M to Accelerate Neocloud Pivot After Nvidia Licensing Deal

Why rare books are so valuable to AI developers

The value of these books lies in their uniqueness. Most text on the internet has already been ingested by AI models, and much of it is low quality or duplicated. Rare books, especially those from before the digital age, offer clean, well-written, and unique content that can help improve an AI’s language understanding and reasoning abilities. Additionally, because these books are out of print, they are not available through other means, giving Amazon a competitive edge in training data.

This is part of a broader trend. Anthropic, for example, was previously sued for using pirated book collections to train its models, and other companies have struck deals with publishers to license their catalogs. Amazon’s approach, however, raises ethical questions about the destruction of cultural artifacts for commercial gain.

Preservationists and authors push back

Librarians, archivists, and authors have expressed concern over the loss of rare books that may not exist in any other copy. While digitization preserves the content, the physical book itself often has historical and monetary value that cannot be replicated. Some rare books are unique editions with annotations, bindings, or illustrations that are lost when the spine is cut.

The practice also highlights a legal gray area. Under the first-sale doctrine, buying a book gives the owner the right to resell or destroy it, but scanning and using the content for commercial AI training may infringe on copyright, especially if the books are still under protection. Amazon has not disclosed whether it obtains permission from rights holders, and the company’s statement did not address the legal implications.

What this means for the AI industry and readers

Amazon’s move is a sign that the AI industry is reaching the limits of available text data. As companies scramble to secure exclusive datasets, the pressure on physical archives and unpublished works will likely increase. For readers and collectors, this could mean that rare books become even scarcer and more expensive, as tech companies buy up inventories.

For AI researchers, the trend raises questions about the sustainability of training data sourcing. If companies resort to destroying physical books, what other sources might they tap next? The answer could shape the future of AI development and the preservation of human knowledge.

Disclaimer: This article is for informational purposes only and does not constitute financial, legal, or investment advice. The cryptocurrency and AI markets are volatile and uncertain; readers should conduct their own research before making any decisions.

CoinPulseHQ Editorial

Written by

CoinPulseHQ Editorial

The CoinPulseHQ Editorial team is a dedicated group of cryptocurrency journalists, market analysts, and blockchain researchers committed to delivering accurate, timely, and comprehensive digital asset coverage. With combined experience spanning over two decades in financial journalism and technology reporting, our editorial staff monitors global cryptocurrency markets around the clock to bring readers breaking news, in-depth analysis, and expert commentary. The team specializes in Bitcoin and Ethereum price analysis, regulatory developments across major jurisdictions, DeFi protocol reviews, NFT market trends, and Web3 innovation.

Be the first to comment

Leave a Reply

Your email address will not be published.


*