PulseAugur
EN
LIVE 14:33:39

Internet Archive eyed for community LLM training data

The Internet Archive is being considered as a potential source for training data for future community-led Large Language Models (LLMs). This approach could address concerns related to data centralization, the environmental impact of AI data centers, and the strain on websites from excessive scraping. A key challenge for such a project would be ensuring that the data used is properly licensed for LLM training, as many websites have restrictive copyright or require attribution. AI

IMPACT Could offer a decentralized alternative for LLM training data, mitigating concerns about resource concentration and environmental impact.

RANK_REASON The item discusses a potential future use case for an existing entity (Internet Archive) in relation to LLM training data, framed as a question and exploring challenges, rather than reporting on a concrete event or release.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Internet Archive eyed for community LLM training data

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Hmm I wonder if the Internet Archive could supply training data for future community # LLM projects? That would seem to solve the concerns with centralization o

    Hmm I wonder if the Internet Archive could supply training data for future community # LLM projects? That would seem to solve the concerns with centralization of wealth/power, # AI data center land/energy/water use, websites getting DDoSed by scrapers, etc. How would such a proje…