Brief · PulseAugur

TOOL · arXiv cs.AI English(EN) · 8h

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff

Researchers have identified a phenomenon called "model plasticity loss" that hinders the effectiveness of Reinforcement Learning (RL) after Supervised Fine-Tuning (SFT) for large language models. Excessive SFT can lead to over-confident token distributions and difficult optimization landscapes, limiting RL's ability to further enhance model capabilities. To address this, a new method called "Rejuvenation" has been proposed, which uses base-anchored model fusion and targeted neuron resets to restore plasticity while retaining SFT benefits, showing improved performance on reasoning and agentic tasks. AI

IMPACT Addresses a key limitation in LLM training pipelines, potentially improving model performance on complex tasks.

Reinforcement Learning
Large Language Model
Supervised Fine-Tuning
Rejuvenation