Back to News Feed
TechCrunch AI15d agoAmanda Silberling

Amazon, which started off selling books, is destroying rare texts to train AI

The High Cost of Data Acquisition

In a startling revelation, it appears that Amazon is systematically dismantling rare literary works to feed the insatiable hunger of its Large Language Model (LLM) training programs. According to an investigation by 404 Media, which utilized a tracking device hidden within a rare volume, these physical books are being funneled into a specialized Amazon facility in Las Vegas known as VGT3. The site, ironically marked by a logo of a dinosaur clutching a book, is reportedly where the spines are removed to facilitate high-speed scanning.

Why Rare Books Matter

As the internet’s accessible data becomes saturated, tech giants are turning to physical archives to sustain AI development. The primary motivations for this aggressive data harvesting include:

  • Data Scarcity: LLMs have already ingested the vast majority of publicly available online text.
  • Authenticity: Older texts provide a "clean" dataset, ensuring the information predates the rise of AI-generated content.
  • Preventing Model Collapse: By avoiding AI-generated training data, developers hope to prevent the degradation of model quality known as "model collapse."

"Amazon purchases books through commercial channels to improve the products and services customers use," the company stated in response to the findings.

While Amazon maintains that its practices are standard for product improvement, the destruction of rare, out-of-print texts highlights the increasingly desperate lengths to which corporations will go to secure high-quality training data for their proprietary models.

#llm#model