AI companies are acquiring large quantities of older books to use as training data, seeking to avoid the "AI slop" of internet-scraped text that could lead to model collapse. Services like ISBNdb facilitate these bulk purchases, offering a way to source pre-2022 printed materials, which are considered structurally free of AI-generated content. This practice has drawn attention due to lawsuits against companies like Anthropic and Google, which were reportedly scanning and destroying books in the process, raising concerns about copyright and the potential loss of these works from society. AI
IMPACT This trend highlights a potential bottleneck in AI training data and may spur new methods for data curation and copyright management.
RANK_REASON The cluster details a significant industry trend of AI companies sourcing physical books for training data, involving major players and raising legal and ethical questions.
- Anthropic
- Better World Books
- Google Gemini
- International Standard Book Number
- The Washington Post
- 404 Media
- Mastodon
- 404media
- AI companies
AI-generated summary · Google Gemini · from 12 sources. How we write summaries →