PulseAugur
EN
LIVE 06:58:50

AI-generated text harms language model pretraining, study finds

A new research paper from arXiv explores the impact of AI-generated text on language model pretraining. The study found that while adding AI-generated tokens can initially lower loss on human text for data-starved models, this benefit quickly reverses into harm as more AI text is added. For models trained on larger datasets, AI tokens immediately increase loss. The researchers propose a new scaling law that accounts for both the benefits and harms of AI text, suggesting that filtering AI text is beneficial when the target is human text and recommending separate validation loss reporting for human and AI-generated text. AI

IMPACT Suggests a need to filter AI-generated content from training datasets to maintain model performance on human text.

RANK_REASON Research paper published on arXiv detailing findings about AI-generated text in pretraining data. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

AI-generated text harms language model pretraining, study finds

How we ranked this

Signal score
25 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Research paper published on arXiv detailing findings about AI-generated text in pretraining data. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi ·

    How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

    arXiv:2609.40295v1 Announce Type: new Abstract: Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31…