A new research paper explores the phenomenon of semantic forgetting in language models, specifically comparing supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT) on classification tasks. The study utilizes a linear-softmax policy to decompose model updates into semantic and stylistic components. It posits that while both methods exhibit parallel semantic updates, SFT can lead to style drift and subsequent forgetting, whereas RFT, under certain conditions, can preserve semantic accuracy. AI
IMPACT Provides theoretical insights into model training dynamics, potentially guiding future fine-tuning strategies to mitigate catastrophic forgetting.
RANK_REASON The cluster contains a single academic paper published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →