PulseAugur
EN
LIVE 18:46:09
ENTITY On-Policy Distillation

On-Policy Distillation

PulseAugur coverage of On-Policy Distillation — every cluster mentioning On-Policy Distillation across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
20
40 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
20
39 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

11 day(s) with sentiment data

RECENT · PAGE 1/2 · 40 TOTAL
  1. TOOL · CL_193974 ·

    New framework FlowErase-OPD enables multi-concept erasure in text-to-image models

    Researchers have developed FlowErase-OPD, a new framework designed to improve safety in text-to-image generation models by enabling the simultaneous erasure of multiple concepts. This method utilizes on-policy distillat…

  2. TOOL · CL_191658 ·

    New distillation technique boosts multilingual math reasoning in LLMs

    Researchers have explored On-Policy Delta Distillation (OPD^2), an advancement over On-Policy Distillation (OPD), for multilingual mathematical reasoning. Experiments using the Qwen3 model demonstrated that OPD^2 signif…

  3. RESEARCH · CL_180528 ·

    Hunyuan3 architecture enhances search agents with cross-domain distillation

    Researchers have developed a new training framework for autonomous search agents called Yuanbao, built on the Hunyuan3 architecture. This framework uses a hybrid approach combining reinforcement learning with cross-doma…

  4. RESEARCH · CL_180525 ·

    New distillation method FTB improves agent performance by validating teacher guidance

    Researchers have developed a new method called FutureBridge-OPD (FTB) to improve on-policy distillation (OPD) for agentic tasks. Standard OPD supervises students on states visited by the teacher, but student deviations …

  5. TOOL · CL_178257 ·

    New MAGA method fuses GUI agents for cross-environment deployment

    Researchers have developed MAGA, a novel method for consolidating specialized GUI agents into a single cross-environment policy. Unlike previous approaches that struggle with conflicting actions or treat all response to…

  6. RESEARCH · CL_180478 ·

    New method RSTG improves LLM reinforcement learning with adaptive teacher guidance

    Researchers have developed RSTG (Recovering Learning Signals via Adaptive Teacher Guidance), a novel method to improve reinforcement learning for large language models. Existing methods like GRPO struggle with sparse re…

  7. TOOL · CL_165044 ·

    New Byte-Prefix Marginalization method improves language model distillation

    Researchers have developed a new method called Byte-Prefix Marginalization (BPM) for on-policy distillation (OPD) of open-weight language models. BPM addresses the challenge of consolidating models with different tokeni…

  8. RESEARCH · CL_160792 ·

    Visual Contrastive Self-Distillation Improves Qwen VL Models

    Researchers have developed Visual Contrastive Self-Distillation (VCSD), a novel method for improving Vision-Language Models (VLMs) without requiring external teachers or privileged information. VCSD works by comparing a…

  9. RESEARCH · CL_163913 ·

    New research enhances LLM inference speed with advanced speculative decoding techniques · 8 sources tracked

    Researchers are exploring advanced techniques to accelerate large language model (LLM) inference through speculative decoding. New methods like "Functional Reconstruction" aim to improve the agreement between draft and …

  10. TOOL · CL_154407 ·

    New research analyzes how SFT, RL, and OPD shape LLM reasoning confidence

    A new research paper introduces a three-stage framework to analyze how supervised fine-tuning (SFT), reinforcement learning (RL), and on-policy distillation (OPD) affect the confidence calibration of large language mode…

  11. RESEARCH · CL_154150 ·

    New benchmarks and methods advance medical vision-language models

    Researchers have developed new benchmarks and distillation techniques to improve the capabilities of vision-language models (VLMs) in the medical domain. PathAgentBench focuses on evaluating VLMs' ability to acquire and…

  12. RESEARCH · CL_156464 ·

    New H$^2$SD framework boosts LLM reasoning via hybrid self-distillation

    Researchers have developed H$^2$SD, a novel hybrid hindsight self-distillation framework designed to enhance the reasoning abilities of large language models. This method addresses limitations in existing reinforcement …

  13. RESEARCH · CL_151917 ·

    New methods enhance LLM post-training with improved RL and data selection

    Researchers have developed new methods to improve large language model (LLM) post-training. Distilled Reinforcement Learning (Distilled RL) integrates teacher supervision into the RL objective to provide fine-grained gu…

  14. RESEARCH · CL_167450 ·

    New frameworks enhance AI model distillation, tackling heterogeneity and spurious signals

    Researchers have developed several new frameworks for on-policy distillation (OPD) to improve AI model capabilities. Any-OPD enables distillation between different model families by using a shared vision representation,…

  15. RESEARCH · CL_147463 ·

    New Contrastive Policy Optimization method improves reinforcement learning

    Researchers have introduced Contrastive Policy Optimization (CPO), a novel method for reinforcement learning with verifiable rewards. CPO utilizes token-level contrastive disagreement between generated text distribution…

  16. TOOL · CL_152483 ·

    New Contrastive Policy Optimization framework enhances reinforcement learning

    Researchers have introduced Contrastive Policy Optimization (CPO), a novel framework for reinforcement learning with verifiable rewards. CPO leverages token-level contrastive disagreement between reference-guided and va…

  17. RESEARCH · CL_152127 ·

    New distillation methods enhance multimodal AI reasoning capabilities

    Researchers have developed new on-policy distillation techniques to improve multimodal AI models. The OPOD method routes student responses to modality-specific teachers, achieving state-of-the-art results across various…

  18. TOOL · CL_145670 ·

    New framework analyzes how LLM training methods affect reasoning confidence

    Researchers have developed a new three-stage framework to analyze how supervised fine-tuning (SFT), reinforcement learning (RL), and on-policy distillation (OPD) affect the confidence of large language models during rea…

  19. RESEARCH · CL_141172 ·

    New research tackles pathologies in On-Policy Distillation for LLMs

    Researchers have identified and proposed solutions for two key pathologies in On-Policy Distillation (OPD), a technique used in large language model post-training. The first pathology, Student-Teacher Mismatch, occurs w…

  20. TOOL · CL_129222 ·

    New H-OPD framework improves multimodal reasoning with dynamic teacher arbitration

    Researchers have introduced H-OPD, a novel framework for multimodal reasoning that enhances on-policy distillation (OPD). Unlike previous methods that use static teacher routing, H-OPD employs a confidence-aware, token-…