AI companies are increasingly sourcing training data from old, printed books, particularly those published before 2022. This strategy aims to avoid "model collapse," a phenomenon where AI models trained on AI-generated text become less effective. Companies like ISBNdb are facilitating these bulk purchases, offering anonymity to AI labs and helping them scan and process millions of books. This trend has been highlighted by lawsuits against Anthropic and Google for allegedly using copyrighted books for training their models. AI
IMPACT This trend highlights a potential bottleneck in AI development and may lead to increased demand for physical books and new methods for data curation.
RANK_REASON The cluster discusses a trend and strategy employed by AI companies rather than a specific product release, research breakthrough, or policy change.
- Anthropic
- Better World Books
- Google Gemini
- International Standard Book Number
- ISBNdb.com
- The Washington Post
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →