Amazon's massive efforts to develop artificial intelligence have been well-documented in recent years. However, a lesser-known aspect of their work is the disposal of rare and valuable books. According to sources close to Amazon, the company has been destroying large collections of rare books for use as training data for its language model, known as LLMs.
These models have already been trained on an enormous amount of online content, including public domain works, academic papers, and even Wikipedia articles. The problem is that many of these sources are no longer available or are restricted by copyright laws. As a result, Amazon's LLMs are being trained on whatever rare books they can find, effectively rendering them obsolete.
The move to destroy rare book collections has raised concerns among preservationists and collectors who believe the loss of these valuable resources is a significant setback for intellectual property rights. While it may seem like a minor issue compared to the benefits of artificial intelligence development, the value of rare books far outweighs the potential drawbacks. Amazon's actions highlight the need for greater consideration and preservation of cultural heritage materials.