A new research paper from arXiv explores the impact of AI-generated text on language model pretraining. The study found that while adding AI-generated tokens can initially lower loss on human text for data-starved models, this benefit quickly reverses into harm as more AI text is added. For models trained on larger datasets, AI tokens immediately increase loss. The researchers propose a new scaling law that accounts for both the benefits and harms of AI text, suggesting that filtering AI text is beneficial when the target is human text and recommending separate validation loss reporting for human and AI-generated text. AI
IMPACT Suggests a need to filter AI-generated content from training datasets to maintain model performance on human text.
RANK_REASON Research paper published on arXiv detailing findings about AI-generated text in pretraining data. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →