A new research paper explores the impact of single training examples on large language models, specifically GPT-2. The study found that while a single passage can be learned and influence predictions shortly after exposure, this effect decays significantly by the end of pre-training. The research utilized 32 GPT-2 models and measured cross-entropy, interpolation loss, and weight displacement to analyze the learning and forgetting process, concluding that such injections relocate the model within its training basin without fundamentally altering its trajectory. AI
IMPACT Investigates the granular influence of training data, potentially informing future data curation and model robustness.
RANK_REASON Research paper published on arXiv detailing experimental findings about LLM training. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →