New methods refine on-policy distillation for enhanced AI capabilities · 7 sources tracked
ByPulseAugur Editorial·[25 sources]·
Researchers are exploring advanced techniques in on-policy distillation (OPD) to enhance language model capabilities by combining multiple "teacher" models. New methods like LEGO-OPD and SAKI focus on factorizing and routing teacher signals to improve visual grounding and reasoning without degrading performance. Other approaches, such as SCOUT and MAESTRO, address the challenge of teacher-student prefix mismatch and optimize teacher intervention strategies to prevent performance degradation and ensure balanced capability integration across different tasks.
AI
IMPACT
These advancements in on-policy distillation could lead to more capable and specialized AI models, improving performance in complex reasoning and multimodal tasks.
RANK_REASON
Multiple research papers introducing novel techniques for on-policy distillation.
arXiv:2610.10460v1 Announce Type: new Abstract: Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillat…
arXiv:2610.10447v1 Announce Type: new Abstract: Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an at…
arXiv cs.CL
TIER_1English(EN)·Yixuan Tang, Yi Yang·
arXiv:2610.09639v1 Announce Type: new Abstract: On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled…
arXiv:2610.02324v1 Announce Type: cross Abstract: Foundation multimodal large language models are designed to support a broad spectrum of capabilities across diverse domains. Multi-teacher on-policy distillation (MOPD) provides an effective framework for consolidating domain-spec…
arXiv cs.AI
TIER_1English(EN)·Xiang Chen, Futao Su, Kong Wang, Jiayi Chen, TanLin Li·
arXiv:2610.02678v1 Announce Type: new Abstract: On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher, but providing such supervision for every rollout requires substantial teacher computation. We introduce Success-Refe…
arXiv cs.LG
TIER_1English(EN)·Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li·
arXiv:2610.02381v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first …
arXiv:2610.02179v1 Announce Type: new Abstract: Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teach…
On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely…
arXiv cs.AI
TIER_1English(EN)·Jaeyun Shin, Hangeol Chang, Jong Chul Ye·
arXiv:2610.00333v1 Announce Type: cross Abstract: Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary …
arXiv:2609.38360v1 Announce Type: cross Abstract: On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundament…
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization…
On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method fo…
arXiv:2609.37510v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student lear…
arXiv:2609.36601v1 Announce Type: new Abstract: On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce S…
Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full …
arXiv:2609.30837v2 Announce Type: cross Abstract: Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on …
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories ar…
On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained …
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillat…
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letti…
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillat…
On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and…
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabel…
arXiv cs.CV
TIER_1English(EN)·Taojie Zhu, Jing Jin, Yuan Xia, Chenyang Ding, Qunshan He, Wanke Xia, Tao Sun, Yan Chen, Jian Wang, Jinjie Gu, Tao Feng·
arXiv:2610.08398v1 Announce Type: new Abstract: On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration. Gradient corrections directly constrain parameter updates under plai…
arXiv:2609.39120v1 Announce Type: new Abstract: On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods primarily focus on enhancing this teacher-side guidance (e.g., by enriching teache…