PulseAugur
中
实时 06:59:26
English(EN) How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

研究发现:AI生成的文本会损害语言模型的预训练

arXiv上的一篇新研究论文探讨了AI生成的文本对语言模型预训练的影响。研究发现,虽然为数据匮乏的模型添加AI生成的Token最初可以降低在人类文本上的损失,但随着更多AI文本的加入,这种好处会迅速转变为损害。对于在更大数据集上训练的模型,AI Token会立即增加损失。研究人员提出了一个新的缩放定律,该定律同时考虑了AI文本的益处和损害,并提出在目标是人类文本时过滤AI文本是有益的,同时建议为人类文本和AI生成的文本分别报告验证损失。 AI

影响 表明需要从训练数据集中过滤掉AI生成的内​​容,以维持模型在人类文本上的性能。

排序理由 在arXiv上发表的研究论文,详细介绍了关于预训练数据中AI生成文本的发现。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现:AI生成的文本会损害语言模型的预训练

本文如何被排名

Signal score
25 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
在arXiv上发表的研究论文,详细介绍了关于预训练数据中AI生成文本的发现。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi ·

    AI代币价值几何?面向野生AI生成网络文本的规模法则

    arXiv:2609.40295v1 Announce Type: new Abstract: Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31…