Researchers are developing new methods for on-policy distillation, a technique used to train smaller AI models by having them learn from the outputs of larger, more capable models. Apple Machine Learning Research has introduced a diagnostic framework to analyze where and why on-policy distillation is effective, finding that the signal is more beneficial on incorrect student rollouts. Meanwhile, new methods like Veto and RG-OPD aim to stabilize training by reformulating the distillation objective or using verifier feedback to filter unreliable teacher signals. Additionally, ReOPD offers an off-environment approach that reuses pre-collected trajectories to speed up multi-turn distillation for agentic tasks. AI
IMPACT These advancements in on-policy distillation could lead to more efficient training of smaller, capable AI models, potentially reducing computational costs and improving performance on complex reasoning and agentic tasks.
RANK_REASON Multiple research papers introducing new methods and analyses for on-policy distillation.
Read on Hugging Face Daily Papers →
- FLNA
- On-Policy Distillation
- Python
- ReOPD
- Replayed-Prefix On-Policy Distillation
- arXiv
- Multi-Turn On-Policy Distillation with Prefix Replay
- reverse-KL distillation
- Reward-Gated On-Policy Distillation
- RG-OPD
- TSD-KD
- Ajay Jaiswal
- Apple Machine Learning Research
- David Harrison
- Fatih İlhan
- Ijun Jang
- Mohammadreza Armandpour
- Veto
AI-generated summary · Google Gemini · from 6 sources. How we write summaries →