Large language models are facing a shortage of training data as the internet is increasingly depleted of usable content. To address this, AI developers are acquiring and destroying vast quantities of physical books, particularly those published after the 1970s, to digitize their content for training. This practice, referred to as "fair use" by some North American legal professionals, involves the physical destruction of books after their data has been extracted, a method that is more legally permissible in the US than in Europe. AI
IMPACT This practice highlights the growing demand for training data and the potential ethical and legal challenges surrounding data acquisition for AI development.
RANK_REASON The item discusses a practice related to AI training data but does not announce a new model, research, or significant industry event.
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →