On-Policy Distillation
PulseAugur coverage of On-Policy Distillation — every cluster mentioning On-Policy Distillation across labs, papers, and developer communities, ranked by signal.
- instance of FLNA 90%
- instance of On-policy self-distillation 70%
- instance of alphaXiv 70%
- instance of Gotit.pub 70%
- competes with Reinforcement Learning with Verifiable Rewards 60%
- other Reinforcement Learning with Verifiable Rewards 60%
- used by Reinforcement Learning with Verifiable Rewards 50%
- competes with Grpo 50%
- other ScienceCast 50%
11 day(s) with sentiment data
-
New framework FlowErase-OPD enables multi-concept erasure in text-to-image models
Researchers have developed FlowErase-OPD, a new framework designed to improve safety in text-to-image generation models by enabling the simultaneous erasure of multiple concepts. This method utilizes on-policy distillat…
-
New distillation technique boosts multilingual math reasoning in LLMs
Researchers have explored On-Policy Delta Distillation (OPD^2), an advancement over On-Policy Distillation (OPD), for multilingual mathematical reasoning. Experiments using the Qwen3 model demonstrated that OPD^2 signif…
-
Hunyuan3 architecture enhances search agents with cross-domain distillation
Researchers have developed a new training framework for autonomous search agents called Yuanbao, built on the Hunyuan3 architecture. This framework uses a hybrid approach combining reinforcement learning with cross-doma…
-
New distillation method FTB improves agent performance by validating teacher guidance
Researchers have developed a new method called FutureBridge-OPD (FTB) to improve on-policy distillation (OPD) for agentic tasks. Standard OPD supervises students on states visited by the teacher, but student deviations …
-
New MAGA method fuses GUI agents for cross-environment deployment
Researchers have developed MAGA, a novel method for consolidating specialized GUI agents into a single cross-environment policy. Unlike previous approaches that struggle with conflicting actions or treat all response to…
-
New method RSTG improves LLM reinforcement learning with adaptive teacher guidance
Researchers have developed RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), a novel method to improve reinforcement learning for large language models. Existing methods like GRPO struggle with sparse re…
-
New Byte-Prefix Marginalization method improves language model distillation
Researchers have developed a new method called Byte-Prefix Marginalization (BPM) for on-policy distillation (OPD) of open-weight language models. BPM addresses the challenge of consolidating models with different tokeni…
-
Visual Contrastive Self-Distillation Improves Qwen VL Models
Researchers have developed Visual Contrastive Self-Distillation (VCSD), a novel method for improving Vision-Language Models (VLMs) without requiring external teachers or privileged information. VCSD works by comparing a…
-
New research enhances LLM inference speed with advanced speculative decoding techniques · 8 sources tracked
Researchers are exploring advanced techniques to accelerate large language model (LLM) inference through speculative decoding. New methods like "Functional Reconstruction" aim to improve the agreement between draft and …
-
New research analyzes how SFT, RL, and OPD shape LLM reasoning confidence
A new research paper introduces a three-stage framework to analyze how supervised fine-tuning (SFT), reinforcement learning (RL), and on-policy distillation (OPD) affect the confidence calibration of large language mode…
-
New benchmarks and methods advance medical vision-language models
Researchers have developed new benchmarks and distillation techniques to improve the capabilities of vision-language models (VLMs) in the medical domain. PathAgentBench focuses on evaluating VLMs' ability to acquire and…
-
New H$^2$SD framework boosts LLM reasoning via hybrid self-distillation
Researchers have developed H$^2$SD, a novel hybrid hindsight self-distillation framework designed to enhance the reasoning abilities of large language models. This method addresses limitations in existing reinforcement …
-
New methods enhance LLM post-training with improved RL and data selection
Researchers have developed new methods to improve large language model (LLM) post-training. Distilled Reinforcement Learning (Distilled RL) integrates teacher supervision into the RL objective to provide fine-grained gu…
-
New frameworks enhance AI model distillation, tackling heterogeneity and spurious signals
Researchers have developed several new frameworks for on-policy distillation (OPD) to improve AI model capabilities. Any-OPD enables distillation between different model families by using a shared vision representation,…
-
New Contrastive Policy Optimization method improves reinforcement learning
Researchers have introduced Contrastive Policy Optimization (CPO), a novel method for reinforcement learning with verifiable rewards. CPO utilizes token-level contrastive disagreement between generated text distribution…
-
New Contrastive Policy Optimization framework enhances reinforcement learning
Researchers have introduced Contrastive Policy Optimization (CPO), a novel framework for reinforcement learning with verifiable rewards. CPO leverages token-level contrastive disagreement between reference-guided and va…
-
New distillation methods enhance multimodal AI reasoning capabilities
Researchers have developed new on-policy distillation techniques to improve multimodal AI models. The OPOD method routes student responses to modality-specific teachers, achieving state-of-the-art results across various…
-
New framework analyzes how LLM training methods affect reasoning confidence
Researchers have developed a new three-stage framework to analyze how supervised fine-tuning (SFT), reinforcement learning (RL), and on-policy distillation (OPD) affect the confidence of large language models during rea…
-
New research tackles pathologies in On-Policy Distillation for LLMs
Researchers have identified and proposed solutions for two key pathologies in On-Policy Distillation (OPD), a technique used in large language model post-training. The first pathology, Student-Teacher Mismatch, occurs w…
-
New H-OPD framework improves multimodal reasoning with dynamic teacher arbitration
Researchers have introduced H-OPD, a novel framework for multimodal reasoning that enhances on-policy distillation (OPD). Unlike previous methods that use static teacher routing, H-OPD employs a confidence-aware, token-…