Technology companies are reportedly purchasing and destroying old books to use their content as training data for AI models. This practice, which involves cutting off book spines and scanning pages, is criticized for its destructive nature and disregard for cultural heritage. While AI requires vast amounts of text, including books, for training, alternative non-destructive scanning methods exist, such as those patented by Google and employed by the Internet Archive. The debate centers on whether digitalization should serve as a public good, preserving both text and the physical artifact, or as a private extraction of data, treating books merely as raw material. AI
IMPACT Raises ethical questions about the long-term preservation of cultural heritage versus the data needs of AI development.
RANK_REASON The article discusses the ethical implications of AI companies destroying books for training data, framing it as a cultural issue rather than a direct AI release or research.
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →