Amazon is purchasing large quantities of rare books, cutting off their spines, and scanning the pages to train artificial intelligence models. Investigative reporting by 404 Media revealed this practice after researchers placed a tracking device inside a rare book. The item ultimately arrived at an Amazon facility located in Las Vegas.
The facility operates under the designation VGT3 and uses a symbol of a dinosaur holding a book in its claws. When questioned by 404 Media, Amazon provided an official statement regarding the activity. The company stated that it purchases books through commercial channels to improve the products and services customers use.
Technology companies require massive volumes of text to train large language models. Many models have already ingested available public internet content. Some firms have also relied on illegally pirated books. Rare items that are out of print or absent from the internet provide a fresh source of coveted training data for these systems.
Materials published before 2022 hold specific value for developers because they predate modern generative artificial intelligence. Large language models face model collapse when they train on too much synthetic text. This phenomenon causes the quality of model outputs to degrade significantly over time.
The reliance on physical texts demonstrates the extreme measures companies take to secure clean training data. As internet sources become depleted, corporate archives and physical bookstores face new demand. Tech firms continue searching for unique material to maintain competitive advantages in artificial intelligence development.
Amazon has not announced any changes to its data collection methods regarding physical books. Observers and researchers will continue monitoring corporate supply chains to track how tech companies acquire rare literary materials for machine learning operations.



