PulseAugur
EN
LIVE 13:45:57

AI firms hoard pre-2022 books for 'clean' training data

AI companies are reportedly acquiring vast quantities of pre-2022 books, scanning and digitizing them. This move is seen as an attempt to secure "clean" data for training AI models, as physical books represent one of the last remaining sources of data not yet saturated with AI-generated content. Anthropic is specifically mentioned as having already processed millions of books, considering it fair use. AI

IMPACT AI companies' aggressive data acquisition strategies may limit access to traditional knowledge sources for future AI development.

RANK_REASON The item is a social media post expressing an opinion about AI data acquisition practices.

Read on Mastodon — sigmoid.social →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

AI firms hoard pre-2022 books for 'clean' training data

COVERAGE [1]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    They poisoned the internet with AI slop. Now they're buying up every pre-2022 book they can find — slicing, scanning, shredding — because print is the last "cle

    They poisoned the internet with AI slop. Now they're buying up every pre-2022 book they can find — slicing, scanning, shredding — because print is the last "clean" data left. Anthropic did it to millions of books and called it fair use. # AI # Anthropic # AIslop # BookTwitter