On-Policy Distillation
PulseAugur coverage of On-Policy Distillation — every cluster mentioning On-Policy Distillation across labs, papers, and developer communities, ranked by signal.
- instance of alphaXiv 90%
- instance of FLNA 90%
- used by Grpo 70%
- competes with Reinforcement Learning with Verifiable Rewards 70%
- used by FLNA 70%
- instance of On-policy self-distillation 70%
- competes with Grpo 70%
- instance of Gotit.pub 70%
- other Reinforcement Learning with Verifiable Rewards 60%
- used by Reinforcement Learning with Verifiable Rewards 50%
- other ScienceCast 50%
10 day(s) with sentiment data
-
New AI distillation method boosts scientific reasoning in models
Researchers have developed a new method called Verifier-Gated Multi-Expert On-Policy Distillation (VG-OPD) to improve the scientific reasoning capabilities of AI models. This technique addresses the limitation of existi…
-
New research reframes diffusion model optimization for reinforcement learning
Researchers have proposed new methods for optimizing diffusion models, particularly in the context of reinforcement learning. One approach, detailed in "Freeze, Share, Shrink," suggests that the action backbone in diffu…
-
New research advances on-policy distillation for LLM training · 6 sources tracked
Researchers are developing advanced techniques for on-policy distillation (OPD), a method used to improve large language models. New approaches like $\gamma$OPD and STRIDE aim to enhance optimization stability and effic…
-
New framework enhances 3D geometry generation with multi-teacher distillation
Researchers have introduced Flow3D-OPD, a novel post-training framework designed to enhance 3D geometry generation models that utilize flow-matching diffusion Transformers. This two-stage approach incorporates multi-tea…
-
TokenRhythm launches NeoHorse-1, an Agent-Native model trained on agent feedback
TokenRhythm, in collaboration with Wuxinqiong, Tsinghua University, Peking University, and Alibaba Group, has introduced NeoHorse-1, an Agent-Native model. This model, available in 4B and 9B versions, aims to integrate …
-
Wang Yunhe's startup releases NeoHorse, an Agent-Native model trained on execution experience
JiYuan LüDòng, a startup founded by former Huawei Noah's Ark Lab director Wang Yunhe, has released its first Agent-Native model, NeoHorse. This model, developed with support from Wuxin Qiong and research contributions f…
-
On-Policy Distillation: Hard Examples Boost LLM Reasoning
A new paper explores the effectiveness of On-Policy Distillation (OPD) for enhancing large language models, particularly focusing on data efficiency and selection. The research found that even a single example (1-shot O…
-
RISE method enhances language model training via self-extrapolation
Researchers have introduced RISE, a novel method for improving language model post-training through self-extrapolating policy distillation. This technique constructs a synthetic teacher from the model's own reinforcemen…
-
New research paper details OPD-then-RL for enhanced LLM reasoning
A new research paper proposes a two-stage approach called OPD-then-RL for improving large language models' reasoning capabilities. This method combines On-Policy Distillation (OPD) with Reinforcement Learning with Verif…
-
New CA-OPD framework improves vision-language models with confidence-aware distillation
Researchers have developed a new framework called Confidence-Aware On-Policy Distillation (CA-OPD) to improve autoregressive vision-language models. This method addresses compounding errors by using teacher confidence t…
-
New distillation methods enhance LLM training and efficiency
Researchers have developed new methods for on-policy distillation (OPD) to improve the training of language models. One approach, Teacher-Gated On-Policy Distillation (TGOPD), verifies teacher reliability at the prompt …
-
Cliff strategy improves LLM reasoning by rewarding first correct steps
Researchers have introduced Cliff, a novel reward shaping strategy for reinforcement learning with verifiable rewards (RLVR) in large language models. Cliff addresses the limitation of existing methods that rely on coar…
-
Cliff method improves LLM reasoning by rewarding correct prefixes
Researchers have introduced Cliff, a novel reward shaping strategy for reinforcement learning with verifiable rewards (RLVR) in large language models. Cliff leverages an off-the-shelf language model to pinpoint the firs…
-
New distillation methods boost LLM training efficiency and accuracy
Researchers have developed new methods for improving the efficiency and accuracy of training smaller language models using distillation techniques. One approach, Teacher-Gated On-Policy Distillation (TGOPD), verifies te…
-
New OPSA method questions on-policy distillation effectiveness
Researchers have questioned the effectiveness of on-policy distillation (OPD) in large language models, finding that its supervision can be noisy and that student models are largely insensitive to this noise. The gains …
-
New research refines on-policy distillation for AI reasoning models
Researchers are exploring on-policy distillation (OPD) for training reasoning models, a technique that uses a teacher model to provide per-token supervision. However, the effectiveness and optimal configuration of OPD r…
-
New method enhances diffusion transformer efficiency for omnimodal generation
Researchers have developed a new method for efficient in-context diffusion transformers that improves omnimodal generation. The technique, called Anchoring Instruction Outside Mask, uses static text anchors to connect v…
-
New Latent-OPD method enhances LMMs for frame-efficient video reasoning
Researchers have introduced Latent-OPD, a novel method for improving the efficiency of Large Multimodal Models (LMMs) in video reasoning. This technique enhances On-Policy Distillation (OPD) by incorporating trajectory-…
-
New framework FlowErase-OPD enables multi-concept erasure in text-to-image models
Researchers have developed FlowErase-OPD, a new framework designed to improve safety in text-to-image generation models by enabling the simultaneous erasure of multiple concepts. This method utilizes on-policy distillat…
-
New distillation technique boosts multilingual math reasoning in LLMs
Researchers have explored On-Policy Delta Distillation (OPD^2), an advancement over On-Policy Distillation (OPD), for multilingual mathematical reasoning. Experiments using the Qwen3 model demonstrated that OPD^2 signif…