From Online Bookstore to AI Consumer: How Amazon Was Caught Shredding Rare Volumes for Machine Learning

By Joe
Freelance Journalist and Editor


Main Facts: The Investigation That Exposed Amazon’s Book-Shredding AI Operation

Today, Amazon stands tall as a global titan of technology, dominating cloud computing infrastructure through Amazon Web Services (AWS), spearheading consumer artificial intelligence developments, and operating a mammoth international e-commerce platform. Yet, despite its sweeping corporate evolution into a multi-trillion-dollar conglomerate, its brand identity remains intrinsically tethered to its humble beginnings as an online bookstore founded in 1994.

As a primary direct seller of physical and digital literature, the proprietor of the ubiquitous Kindle e-reader line introduced in 2007, and the host of a massive marketplace utilized by countless independent third-party booksellers, Amazon has long styled itself as a champion of the written word—even famously adopting the moniker "Earth’s biggest bookstore."

That carefully curated legacy has suffered a profound and ironic blow. A recent, meticulous investigative report by 404 Media has revealed that Amazon has been systematically tearing apart old, rare, and irreplaceable physical books to feed data directly into its proprietary large language models and artificial intelligence systems.

The revelation lays bare a startling corporate reality: the enterprise built on delivering books to doorstep readers is now consuming them on an industrial scale. To make the optics even more damaging, warehouse workers assigned to this specialized scanning operation reportedly operate under an internal, tone-deaf cartoon logo depicting a Tyrannosaurus rex devouring a book.

The investigation began when a vigilant independent bookseller on Biblio—an online marketplace where buyers can complete transactions anonymously—grew suspicious of a bulk order comprising 1,000 volumes. Suspecting that the buyer might be an artificial intelligence firm harvesting text, the vendor embedded an Apple AirTag inside the pages of one of the books before dispatching the shipment.

The tracking device didn’t lead to a local library, a university archive, or a traditional recycling plant. Instead, the signal traveled straight to Amazon’s sprawling LAS8 warehouse facility in Las Vegas, Nevada—the confirmed operational hub for a specialized scanning unit known internally as VGT3.


Chronology: Tracking the Shipment and Uncovering the VGT3 Scanning Facility

The path from a quiet independent online bookstore to an industrial-scale AI data-harvesting facility unfolded over a series of deliberate operational steps, exposed piece by piece through investigative journalism and whistleblower accounts.

  • The Suspicious Order: An anonymous buyer places a bulk order for 1,000 books via Biblio, prompting immediate skepticism from the seller regarding the true intended use of the literature.
  • The Surveillance Trap: Anticipating potential data-harvesting practices, the bookseller conceals an Apple AirTag inside one of the volumes within the bulk shipment and tracks its transit across the country.
  • The Destination Unveiled: The tracking signal terminates at Amazon’s LAS8 fulfillment and logistics center in Las Vegas, directly pointing to a covert scanning operation designated as VGT3.
  • The Destruction Protocol: According to firsthand accounts from workers inside the facility, staff members assigned to the VGT3 unit systematically slice the physical bindings off books to flatten pages and accelerate the high-speed scanning process required for AI ingestion.
  • The Internal Reveal: Employees on digital worker forums confirm the daily operations of the unit, noting a rigid division of labor where some workers receive and scan barcodes, while others are explicitly tasked with destroying the physical integrity of the volumes.
  • Public Exposure: 404 Media publishes its findings, detailing the AirTag tracking, the physical destruction of rare print works, and the existence of a satirical T-rex corporate logo celebrating the consumption of literature.

Supporting Data: Why AI Companies Target Old and Rare Print

The disclosure of Amazon’s book-shredding operations shines a harsh light on the hidden mechanics of modern artificial intelligence development. As tech giants race to secure high-quality training datasets, the raw material required to train generative AI models has become fiercely contested.

Industry insiders and independent analysts point out that pre-2022 texts hold immense value for AI developers. Modern internet-scraped text is increasingly polluted by machine-generated content, spam, and synthetic writing, which can degrade the performance and accuracy of subsequent AI models if used for training. In contrast, older, physically printed volumes offer "clean," human-authored prose that is uncontaminated by the digital noise of contemporary web content.

Furthermore, rare and antique books often contain specialized historical, cultural, and niche data that competitors cannot easily access through standard digital archives or public domain repositories. While these acquired books may hold modest monetary value on secondary marketplaces, the bookseller who tracked the shipment emphasized that their true worth lies elsewhere:

"There are a lot of other types of value," the vendor noted. "There’s historical value, intellectual value, sentimental value. All sorts of things, and all of those the AI companies don’t care about. They just want the content as a bunch of words strung together."

Amazon's wanton destruction of rare books is a betrayal of the brand's origins

The practice of destroying physical media to feed artificial intelligence bears an eerie resemblance to recent cultural controversies in the tech sector. A couple of years ago, Apple faced a fierce public backlash that forced the company to hastily withdraw an iPad Pro commercial depicting a hydraulic press violently crushing physical creative tools, including musical instruments, paint cans, books, and camera lenses. While Apple’s imagery was purely symbolic and metaphorical, Amazon’s industrial destruction of rare literature represents a literal, mechanical crushing of cultural heritage for corporate gain.


Official Responses and Industry Context

Faced with mounting scrutiny from cultural preservationists, authors, and everyday readers, Amazon has offered a measured corporate response defending its procurement methods.

In an official statement addressing the investigation, an Amazon spokesperson stated:

"Amazon purchases books through commercial channels to help develop and improve the products and services our customers use."

The company has not explicitly denied the existence of the VGT3 scanning facility or the practice of unbinding physical books to accelerate data processing. From a strict legal standpoint, legal experts note that training artificial intelligence models on legally purchased books can sometimes fall under the legal doctrine of fair use in certain jurisdictions. However, legal compliance does little to soothe the public relations fallout.

The strategy adopted by Amazon also stands in stark contrast to the policies of other artificial intelligence developers. While rival firms have similarly faced criticism for acquiring large volumes of literature to train their models, some competitors have drawn clear ethical lines. For instance, AI firm Anthropic has publicly claimed that it explicitly avoids utilizing rare or antique books in its training pipelines, making Amazon’s aggressive, destructive approach appear particularly reckless to cultural historians and rare-book collectors.

Alternative methods to achieve the same technological ends—such as non-destructive book scanning technologies or the licensing of established digital library archives—are widely available within the industry. Choosing instead to tear apart physical artifacts highlights a prioritization of processing speed and cost efficiency over the preservation of tangible history.


Implications: Cultural Vandalism, Brand Alienation, and the Future of AI Ethics

The revelation that Amazon is shredding rare and irreplaceable books to train artificial intelligence models carries sweeping implications for the tech industry, the literary community, and the e-commerce giant’s public standing.

1. Alienation of Authors and Bibliophiles

Amazon owes much of its historical legitimacy and foundational success to the writing and publishing community. By engaging in the physical destruction of literature—regardless of whether the individual titles possessed high monetary value—the company risks alienating the very authors, publishers, bibliophiles, and loyal consumers who view books not merely as data points, but as cultural artifacts worthy of preservation. The dissonance between Amazon’s marketing as a haven for readers and its internal operations as a consumer of physical books creates a profound trust deficit.

2. Deepening Skepticism Toward AI Claims

For years, major technology companies have pitched artificial intelligence as a revolutionary tool destined to elevate human creativity, preserve knowledge, and assist cultural endeavors. Stories like the VGT3 warehouse operation undermine those high-minded marketing claims. When AI development relies on the literal dismantling and disposal of human cultural heritage, it reinforces the growing public perception that corporate AI expansion operates at the direct expense of human artistic and literary legacy.

3. Regulatory and Ethical Precedents

As copyright lawsuits, creator strikes, and intellectual property disputes continue to shadow the generative AI boom, incidents like this add combustible material to the ongoing debate over how training data is sourced. Even if courts ultimately rule that purchasing and scanning physical books satisfies narrow statutory definitions of copyright compliance, the ethical boundary has clearly shifted. Stakeholders, lawmakers, and consumer advocacy groups will likely demand greater transparency from tech conglomerates regarding how their foundational models are built, forcing companies to account not just for what data they use, but how they acquire and treat the physical vessels that carried human knowledge across generations.

Leave a Reply

Your email address will not be published. Required fields are marked *