Researchers have identified and proposed solutions for two key pathologies in On-Policy Distillation (OPD), a technique used in large language model post-training. The first pathology, Student-Teacher Mismatch, occurs when a significant gap between the teacher and student models leads to misaligned guidance. The second, Length Exploitation, arises when the model learns to manipulate response length for higher rewards. To address these issues, new methods like advantage clipping, log-scale compression, and Adaptive Dual-Perspective OPSD (AD-OPSD) are introduced to regulate the distillation signal and preserve the model's native reasoning capabilities, showing improved accuracy on benchmarks. AI
IMPACT These findings could lead to more stable and effective LLM training, improving performance on complex reasoning tasks.
RANK_REASON The cluster contains two arXiv papers detailing research into LLM training methodologies and proposing new techniques.
- Adaptive Dual-Perspective OPSD
- AD-OPSD
- arXiv
- Hugging Face
- large-language models
- On-policy self-distillation
- Thinking Collapse
- alphaXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Length Exploitation
- LLM
- On-Policy Distillation
- RLVR
- ScienceCast
- Student-Teacher Mismatch
AI-generated summary · Google Gemini · from 6 sources. How we write summaries →