Researchers have questioned the effectiveness of on-policy distillation (OPD) in large language models, finding that its supervision can be noisy and that student models are largely insensitive to this noise. The gains from OPD appear to stem from suppressing low-probability tokens rather than direct teacher guidance. This has led to the development of On-Policy Self-Adaptation (OPSA), a new supervision-free method that uses entropy-adaptive negative advantages to improve model performance. OPSA significantly boosts reasoning capabilities, outperforming OPD on benchmarks like AIME24. AI
IMPACT Proposes a more effective, supervision-free method for improving LLM reasoning, potentially reducing reliance on noisy teacher models.
RANK_REASON Academic paper proposing a new method and analyzing existing ones.
Read on Hugging Face Daily Papers →
- AIME24
- arXiv
- Hugging Face
- On-Policy Distillation
- On-Policy Self-Adaptation
- Qwen3-1.7B
- Qwen3-4B-OPSA
- Reinforcement Learning with Verifiable Rewards
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →