Researchers have identified that catastrophic forgetting in large language models, a phenomenon where models lose previously learned information during continued training, is localized in the output embeddings of tokens not present in the new training data. This forgetting is driven by persistent, one-sided softmax gradients amplified by Adam's second-moment normalization. To combat this, a proposed solution involves increasing Adam's epsilon specifically for the output projection layer during training. This adjustment has shown significant reductions in forgetting across various model sizes and families without negatively impacting new learning or requiring model-specific tuning, and it complements existing replay-based methods. AI
IMPACT This research offers a potential method to significantly reduce catastrophic forgetting in LLMs, improving their ability to learn continuously without losing prior knowledge.
RANK_REASON The cluster contains a research paper detailing a novel finding and proposed solution for a specific problem in LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →