PulseAugur
实时 14:13:29
English(EN) Hmm I wonder if the Internet Archive could supply training data for future community # LLM projects? That would seem to solve the concerns with centralization o

互联网档案馆被视为社区LLM训练数据来源

互联网档案馆正被考虑作为未来社区主导的大型语言模型(LLMs)的潜在训练数据来源。这种方法可以解决数据中心化、AI数据中心的环保影响以及过度抓取对网站造成的压力等问题。此类项目的一个关键挑战将是确保所用数据已获得LLM训练的适当许可,因为许多网站都有严格的版权规定或要求署名。 AI

影响 可能为LLM训练数据提供去中心化的替代方案,缓解资源集中和环境影响的担忧。

排序理由 该条目讨论了一个现有实体(互联网档案馆)在LLM训练数据方面的潜在未来用例,以问题的形式提出并探讨了挑战,而不是报道具体的事件或发布。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

互联网档案馆被视为社区LLM训练数据来源

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    嗯,我想知道互联网档案馆是否能为未来的社区#LLM项目提供训练数据?这似乎能解决中心化的问题

    Hmm I wonder if the Internet Archive could supply training data for future community # LLM projects? That would seem to solve the concerns with centralization of wealth/power, # AI data center land/energy/water use, websites getting DDoSed by scrapers, etc. How would such a proje…