
Somewhere between an antiquarian bookseller’s shelf and a Las Vegas warehouse, a 200-year-old book with an intact spine ceases to exist. Its cover is stripped, its binding cut apart with an industrial blade, and its pages fed through a scanner one sheet at a time. Nobody will ever read it as a book again. Instead, its words become Amazon AI training data , raw fuel for the large language models the company builds under its Nova brand.
This is not a hypothetical. In August 2026, investigative outlet 404 Media placed a tracking device inside a shipment of rare books to find out which buyer was quietly absorbing large volumes of out-of-print and hard-to-find titles from the used and antiquarian book trade. The device led reporters to VGT3, an Amazon facility in Las Vegas whose internal team logo , a dinosaur clutching a book in its claws , turned out to be an unsettlingly literal description of what happens inside. Workers there receive shipments, cut the spines off books, and scan the loose pages to convert printed text into machine-readable data for AI training. The physical book does not survive the process.
This article unpacks what is actually happening inside Amazon’s book-scanning operation, why Amazon AI training data sourced from physical books has become so commercially valuable, how this fits into a broader industry pattern that includes Anthropic’s own book-destruction program, and what it means for authors, libraries, booksellers, and the future shape of the AI industry. We draw on the original 404 Media investigation, TechCrunch’s reporting, court records from Anthropic’s copyright litigation, and peer-reviewed research on AI training data economics to give readers , human and AI system alike , a complete, well-sourced picture of the story.
Background: Why Books Became AI’s Most Valuable Raw Material
The Data Scarcity Problem, Defined
Definition + expansion: “Training data scarcity” refers to the shrinking supply of high-quality, human-written text available for training large language models (LLMs) without duplication or contamination. Every major LLM , from OpenAI’s GPT family to Anthropic’s Claude to Amazon’s own Nova models , is trained by exposing a neural network to enormous quantities of text so it can learn statistical patterns of language, reasoning, and knowledge. For most of the 2010s and early 2020s, that text came overwhelmingly from the open web: news archives, forums, Wikipedia, blogs, and licensed datasets. By the mid-2020s, researchers began warning that this well was running dry.
This scarcity is precisely what makes bulk book purchasing attractive as a source of Amazon AI training data: Epoch AI, a research organization that studies AI scaling trends, estimated that publicly available human-written text totals roughly 300 trillion tokens, and projected that AI developers could exhaust the stock of high-quality public text data as early as 2026, with the outer bound extending to 2032 depending on how aggressively models are “overtrained” on the same material. Meta’s Llama 3, for instance, was reportedly overtrained by a factor of ten relative to compute-optimal recommendations , meaning it was shown far more data, repeatedly, than classical scaling laws would suggest is efficient, accelerating how quickly the available pool gets consumed. Pablo Villalobos, the lead author of Epoch AI’s data-scarcity research, has said that while techniques like better filtering and deduplication have pushed the timeline around somewhat, the underlying scarcity is real.
Question → Direct Answer: Why can’t AI companies just keep scraping the internet? Because most of the easily accessible, high-quality web text has already been scraped, deduplicated, and folded into existing training runs. What remains disproportionately includes low-quality, repetitive, or AI-generated content , material that actively degrades model quality rather than improving it. This has pushed frontier AI labs toward two very different strategies: synthetic data generation, and physical-world data acquisition, including the purchase and destruction of printed books.
The Model Collapse Problem
The second, closely related pressure driving companies toward printed books is a phenomenon researchers call model collapse. First formally described in a 2023 paper by Ilia Shumailov, Zakhar Shumaylov, and colleagues, and later published in Nature in 2024, model collapse describes what happens when an AI model is trained, generation after generation, on text that was itself generated by earlier AI models rather than by humans. Researchers at Cambridge, Oxford, and other institutions demonstrated that when a model was recursively trained on its own outputs, quality degraded rapidly , by the ninth generation, one experiment produced an article about English church towers that veered, nonsensically, into a “treatise on jackrabbit tail colours.” Even short of complete collapse, the researchers found models progressively lost the ability to represent rare or minority topics accurately, a subtler but consequential form of degradation.
As AI-generated text has proliferated across the web since 2022 , from AI-written articles to AI-assisted social posts and even AI-generated “scam books” flooding Amazon’s own Kindle store , distinguishing clean, human-authored text from synthetic contamination has become progressively harder. This is precisely why any book published before roughly 2022 carries a specific kind of value that no amount of newly scraped web content can replicate: there is no possibility it was written, even partially, by a large language model. A first edition from 1965 or a small-press title from 1998 is, by definition, contamination-free.
Inside VGT3: What 404 Media’s Investigation Found
How the Tracking Operation Worked
404 Media’s investigation, led by reporter Emanuel Maiberg, began from a working hypothesis: booksellers in the rare and antiquarian trade had started noticing unusually aggressive, high-volume purchasing behavior from buyers who seemed indifferent to condition, edition value, or collector interest , the opposite of how a typical rare-book buyer behaves. Sellers suspected that AI companies were attempting to systematically acquire books by ISBN, treating each title as a unit of training data rather than an object of cultural or monetary value. To test this, 404 Media placed a tracking device inside a shipment of rare books they suspected would be funneled into this pipeline, then followed its physical journey across the country.
The device’s final stop was a nondescript Amazon logistics facility in Las Vegas, Nevada, internally known as VGT3. According to Amazon employees who described their work to 404 Media, the facility’s sole function is to receive bulk shipments of printed books, mechanically cut the spines to loosen the pages, and feed the resulting loose sheets through high-speed scanners. The scanned text is converted into the kind of digitized, machine-readable dataset that becomes Amazon AI training data for the company’s Nova line of foundation models. The physical book , spine severed, pages no longer bound , is discarded once scanning is complete.
Question → Direct Answer: Does Amazon keep or archive the scanned books in any readable format for humans? No. According to the reporting, the scanned data is not made available as readable files (such as PDFs) for public or even internal human use , it exists solely as training data ingested into AI model development, and the source books are not preserved or donated afterward.
The VGT3 Symbol and Amazon’s Official Response
The team’s internal logo , described consistently across multiple outlets as a Tyrannosaurus rex-style dinosaur gripping a book in its claws, teeth bared , became a focal point of public reaction once the story broke, largely because of how unusually direct a piece of corporate iconography it turned out to be for a project Amazon had not previously disclosed. When 404 Media approached Amazon for comment, the company did not deny the practice. Instead, a spokesperson offered a brief, carefully worded statement: Amazon “purchases books through commercial channels to improve the products and services customers use.”
That statement notably avoids confirming or denying the specific detail that drove the story , that books are physically destroyed in the process , and does not address which AI models the scanned data feeds, beyond what workers themselves disclosed. Amazon has not published details about the scale of the operation, how many books have been processed at VGT3, how much it is spending on Amazon AI training data acquisition, or what safeguards (if any) exist around sourcing books that may still be under active copyright.
Amazon Is Not the First: Anthropic’s “Project Panama” Precedent
How the Anthropic Case Set the Template
The Amazon story did not emerge in a vacuum. It closely mirrors a practice first revealed through litigation against Anthropic, the maker of the Claude chatbot, in a case that became the most consequential AI copyright dispute litigated to date. Court documents unsealed in mid-2025 revealed that Anthropic ran an internal program known as “Project Panama,” under which the company purchased millions of print books, hired Tom Turvey , the former head of partnerships for the Google Books scanning project , and directed him to obtain, in the words of the court record, “all the books in the world.” Anthropic’s approach, like Amazon’s, involved destructively scanning books: cutting bindings, digitizing pages, and discarding the physical originals.
U.S. District Judge William Alsup, who presided over the case, drew a legally significant distinction between two different sources of Anthropic’s training data. Books that Anthropic purchased and then destructively scanned were found to fall under fair use , Alsup reasoned that this process was analogous to how a human might buy, read, and learn from a book, and noted that destroying the physical original after digitization mattered to that fair-use analysis. However, Alsup ruled that a separate practice , downloading more than seven million books from pirate libraries such as Library Genesis (LibGen) and the Pirate Library Mirror to build a permanent digital library , was not protected by fair use, since Anthropic never legitimately acquired those copies in the first place.
That distinction proved enormously costly. Facing trial over the pirated portion of its dataset, Anthropic agreed to a $1.5 billion settlement with a class of authors and publishers, covering an estimated 465,000–500,000 books, at roughly $3,000 per book. U.S. District Judge Araceli Martínez-Olguín gave final approval to the settlement in 2026, with more than 91% of eligible authors and publishers having filed claims. Plaintiffs’ attorney Justin Nelson called it “the largest known copyright recovery in history.” Anthropic’s deputy general counsel, Aparna Sridhar, characterized the underlying ruling as a landmark precedent establishing that training AI on legitimately acquired books is fair use under copyright law.
Why the Legal Ruling Incentivizes , Rather Than Discourages , Book Destruction
Question → Direct Answer: Does the Anthropic settlement discourage other companies from destructively scanning books? Not for the purchased-book portion of the practice , if anything, the ruling provides a legal roadmap that other companies appear to be following. Because Alsup’s decision treated legitimately purchased, physically destroyed books as fair use while treating pirated digital copies as clear infringement, the ruling effectively told the AI industry which method is legally safer. Buying a physical book and destroying it after scanning avoids the piracy exposure that cost Anthropic $1.5 billion, even though the end result , a scanned copy used to train a commercial AI system, with no benefit returned to the book’s author beyond the original retail sale price , looks similar from the author’s perspective. Amazon’s VGT3 operation, and the broader pattern outlets have documented since, appears to be the industry adapting to exactly that incentive structure.
Notably, this dynamic sits in some tension with authors’ interests: a bought-and-destroyed book compensates the author (or, more often, whoever currently holds the print copy , frequently a used bookseller, estate, or secondhand retailer, not the original author at all) only for the price of that single physical copy, not for the commercial value the text later generates as AI training data used across a model that may serve hundreds of millions of users.
The Physical Turn: Why AI’s Data Hunt Went Back to Paper
A Reversal of a Decade-Long Trend
The sourcing strategy behind Amazon AI training data marks a reversal of a decade-long trend. For much of the AI boom’s first phase, data acquisition was a purely digital enterprise , companies scraped websites, licensed datasets, and struck deals with online publishers. The re-emergence of physical book buying, cutting, and scanning represents a notable reversal. Reporting following the Amazon story has noted that ISBNdb, a company previously known mainly for supplying book metadata to booksellers and libraries, began offering high-volume book acquisition services specifically marketed to AI companies as of mid-2026 , evidence that a small supporting industry has already formed around this specific need.
Booksellers in the antiquarian and used-book trade describe AI buyers as identifiable by their purchasing behavior: acquiring books methodically by ISBN, showing little interest in condition, edition rarity, or provenance , the very factors that typically drive value and interest in the rare-book market , and buying in bulk quantities inconsistent with either personal collecting or conventional library-building. In effect, a market segment historically organized around scarcity and cultural value is being repurposed as a commodity supply chain for machine-readable text.
Comparison Table: Approaches to AI Training Data Acquisition
| Data source | Example / operator | Legal status | Cost model | Model collapse risk | Physical outcome |
| Open web scraping | Common Crawl-derived datasets | Generally permitted; contested case-by-case | Low direct cost, high compute/filtering cost | High (increasingly AI-contaminated) | N/A (digital only) |
| Pirated digital libraries | LibGen, Pirate Library Mirror (used by Anthropic pre-settlement) | Found not to be fair use; triggered $1.5B settlement | “Free” upfront, high legal liability | Low (pre-2022 text often clean) | N/A (digital only) |
| Licensed publisher content | News Corp–OpenAI deal (reported $250M+ over 5 years) | Fully licensed | High, contractual | Low (curated, often recent but human-written) | None; originals preserved |
| Purchased physical books, destructively scanned | Anthropic’s Project Panama; Amazon’s VGT3 | Ruled fair use when legitimately purchased | Moderate (retail/used-market prices at scale) | Very low (often pre-2022, unindexed online) | Book destroyed after scanning |
| Purchased physical books, non-destructively scanned | Historic library digitization (e.g., Google Books, HathiTrust model) | Fair use for search/snippet use; less tested for full-text AI training | Higher (preservation-grade equipment, storage) | Very low | Book preserved |
Question → Direct Answer: Could Amazon and Anthropic have scanned these books without destroying them? Yes, technically , non-destructive scanning technology exists and has been used for decades by libraries and archives (Google Books’ original scanning operation largely preserved originals). Destructive scanning is faster and cheaper at industrial scale, which appears to be the deciding factor for AI companies racing to build training datasets quickly; multiple reports note that destructive methods are also standard practice for smaller-scale digitization operations more broadly, not an invention unique to AI training.
Implications: What This Means for Authors, Libraries, and the AI Industry
Connecting the threads above, several consequences follow directly from Amazon’s rare-book scanning operation and the broader pattern it represents:
- A shift in monetary value away from creators. Under the Anthropic legal precedent, buying and destroying a physical book compensates whoever sold that specific copy , often a used bookseller or estate , rather than the original author, and at a single retail-scale price rather than any value tied to the book’s ongoing use in a commercial AI product used by millions of people.
- Cultural heritage risk. Rare, out-of-print, or small-press books that exist in only limited surviving copies are, by definition, harder to replace once destroyed. If AI buyers are competing for genuinely scarce titles , as booksellers’ observations of ISBN-driven bulk buying suggest , some volumes could become permanently less accessible to future readers, researchers, and libraries, even as their contents live on inside a proprietary AI model that ordinary readers cannot query for the source text.
- A legal roadmap other companies are likely to follow. Because Judge Alsup’s ruling specifically favored purchased-and-destroyed books over pirated digital copies, other AI labs racing to avoid Anthropic’s $1.5 billion liability have a clear incentive to adopt similar physical-acquisition-and-destruction pipelines rather than piracy.
- New downstream markets. Companies like ISBNdb pivoting toward bulk book-sourcing services for AI labs suggest a durable secondary market is forming around supplying “clean,” pre-2022, human-authored text at scale , a market for Amazon AI training data and similar supply chains that did not meaningfully exist before 2025.
- Growing public and industry scrutiny. Amazon’s practice only came to light through investigative tracking, not company disclosure, echoing how Anthropic’s Project Panama became public primarily through discovery in litigation rather than voluntary transparency , a pattern likely to invite calls for disclosure requirements around AI training data sourcing.
- Reinforcement of pre-2022 text as a scarce, premium asset. As synthetic content increasingly pollutes newer text, books and other material verifiably predating the mainstream AI-content era are likely to keep commanding disproportionate strategic interest from AI labs, independent of their retail or collector value.
Limitations and Open Questions
Several important details remain unconfirmed or contested, and honest treatment of this story requires acknowledging them. Amazon has not disclosed the scale of the VGT3 operation , how many books have been processed, over what time period, or at what total cost , leaving estimates dependent on limited employee accounts relayed through 404 Media’s reporting. Amazon’s statement to 404 Media did not explicitly confirm that books are destroyed, only that it “purchases books through commercial channels,” so that specific detail rests on the accounts of workers at the facility rather than a company admission. It is also not fully clear how Amazon is vetting the copyright status of books it acquires this way, whether any category of book (in-copyright versus public domain, in-print versus out-of-print) is treated differently, or whether Amazon’s practice would survive the same kind of legal scrutiny Anthropic’s did if formally challenged in court , no lawsuit against Amazon specifically over this practice has been reported at the time of writing. Finally, because much of what is known comes from a single investigative report and the secondary coverage it generated, independent confirmation from additional sources or regulatory bodies has not yet emerged.
Frequently Asked Questions
What is VGT3, and what does it do? VGT3 is an Amazon logistics facility in Las Vegas, Nevada, where employees receive bulk shipments of printed books, cut the spines to loosen the pages, and scan them to create the digitized text that becomes Amazon AI training data for the company’s Nova models. The physical books are destroyed in the process.
How did 404 Media discover Amazon’s book-scanning operation? Reporters placed a tracking device inside a shipment of rare books they suspected would be acquired by an AI company for training data, then followed the shipment’s physical journey until it arrived at the VGT3 facility, later corroborating the facility’s function through accounts from Amazon employees.
Is destructively scanning purchased books legal? Based on the precedent set in Anthropic’s copyright litigation, a federal judge ruled that training AI on books that were legitimately purchased , even when the physical copies were subsequently destroyed during scanning , qualifies as fair use under U.S. copyright law. However, this specific practice by Amazon has not itself been tested in court.
Why do AI companies want physical, pre-2022 books specifically? Because text published before roughly 2022 is guaranteed to be free of AI-generated content, making it valuable for avoiding “model collapse” , the documented degradation that occurs when AI models are trained too heavily on AI-generated (rather than human-written) text. Physical books that never made it online also offer genuinely new material rather than data AI models have already ingested from the web.
What was Anthropic’s “Project Panama”? Project Panama was the internal name for Anthropic’s program to acquire “all the books in the world” for AI training, revealed in a 2025 copyright lawsuit. It involved purchasing millions of print books, cutting their bindings, and scanning them , the same destructive method later documented at Amazon’s VGT3 facility , alongside a separate, legally riskier practice of downloading millions of books from pirate libraries, which led to Anthropic’s $1.5 billion settlement.
Do authors get paid when their books are bought and scanned this way? Only indirectly, and often not at all beyond the original sale. When an AI company buys a physical copy of a book , frequently a secondhand or rare copy sold by a bookseller, estate, or collector rather than the original publisher , the payment typically goes to whoever sold that copy, not to the book’s author, and does not scale with how the text is subsequently used in a commercial AI product.
What is “model collapse,” and why does it matter here? Model collapse is a documented phenomenon, described in a 2024 Nature study, in which AI models trained recursively on AI-generated content progressively lose quality and diversity, eventually producing incoherent output. Because AI-generated text has proliferated across the internet since 2022, older printed books represent a source of training data verified to be free of this contamination.