On-policy self-distillation
PulseAugur coverage of On-policy self-distillation — every cluster mentioning On-policy self-distillation across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
New SCOPE-OPSD method improves AI model self-distillation
Researchers have developed a new method called SCOPE-OPSD, which enhances on-policy self-distillation (OPSD) by incorporating Fisher-conditioned privileged subspaces. This technique aims to improve the transfer of super…
-
VISTA method enhances AI reasoning via teacher-student adaptation · 2 sources tracked
Researchers have developed VISTA, a novel method for on-policy self-distillation (OPSD) that enhances AI model reasoning. Unlike standard OPSD, VISTA adapts the teacher model based on verified student rollouts, particul…
-
New TTPO method enhances LLM math reasoning without labels
Researchers have developed Test-Time Policy Optimization (TTPO), a novel method for improving large language models' mathematical reasoning capabilities without relying on ground-truth labels. TTPO addresses the fragili…
-
New TTPO method enhances LLM reasoning without labels
Researchers have developed Test-Time Policy Optimization (TTPO), a novel method for improving large language models' mathematical reasoning capabilities without relying on ground-truth labels. TTPO addresses the fragili…
-
New research refines on-policy distillation for AI reasoning models
Researchers are exploring on-policy distillation (OPD) for training reasoning models, a technique that uses a teacher model to provide per-token supervision. However, the effectiveness and optimal configuration of OPD r…
-
ArmorOCR framework enhances adversarial OCR perception with new AdvSpot benchmark
Researchers have introduced ArmorOCR, a novel two-stage training framework designed to enhance the robustness of optical character recognition (OCR) against adversarial attacks. This framework addresses the limitations …
-
New framework unifies on-policy self-distillation for LLM reasoning · 3 sources tracked
Researchers have developed a unified framework for on-policy self-distillation (OPSD) to enhance LLM reasoning by integrating privileged information into model parameters. This new framework, Unified On-Policy Self-Dist…
-
New AI training method uses rubrics as privileged information for open-ended generation
Researchers have developed a new method called On-policy self-distillation (OPSD) that utilizes rubrics as privileged information (PI) for open-ended text generation. This approach enhances the training signal for model…
-
New method CSCR improves LLM long-context reasoning by reallocating token credit
Researchers have developed a new method called Counterfactual Sensitivity Credit Reallocation (CSCR) to improve the reasoning capabilities of large language models, particularly in tasks requiring long-context reasoning…
-
Visual Contrastive Self-Distillation Improves Qwen VL Models
Researchers have developed Visual Contrastive Self-Distillation (VCSD), a novel method for improving Vision-Language Models (VLMs) without requiring external teachers or privileged information. VCSD works by comparing a…
-
New distillation method trains LLMs efficiently with soft prompts
Researchers have developed a new method called Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context ("method") to train large language models. This technique uses a teacher model that differs from the st…
-
New research identifies decoding collapse in AI agent self-distillation
Researchers have identified a failure mode in feedback-augmented self-distillation for retrieval-interleaved search agents, termed decoding collapse. This occurs when models generate diverse-looking but input-agnostic r…
-
New H$^2$SD framework boosts LLM reasoning via hybrid self-distillation
Researchers have developed H$^2$SD, a novel hybrid hindsight self-distillation framework designed to enhance the reasoning abilities of large language models. This method addresses limitations in existing reinforcement …
-
New on-policy distillation methods enhance LLM reasoning and efficiency · 10 sources tracked
Multiple research papers explore advancements in on-policy distillation (OPD) techniques for language models, aiming to improve reasoning capabilities and training efficiency. Several methods, including SimpleOPD, S$^2$…
-
New research tackles pathologies in On-Policy Distillation for LLMs
Researchers have identified and proposed solutions for two key pathologies in On-Policy Distillation (OPD), a technique used in large language model post-training. The first pathology, Student-Teacher Mismatch, occurs w…
-
New research challenges on-policy self-distillation for LLMs, proposing refined methods · 10 sources tracked
Recent research papers explore the limitations and potential improvements of on-policy self-distillation (OPSD) for training large language models (LLMs). Studies indicate that standard OPSD can lead to rote memorizatio…
-
New methods enhance VLM accuracy for GUI grounding tasks · 2 papers
Two new research papers introduce novel methods for improving the accuracy and reliability of vision-language models (VLMs) in GUI grounding tasks. The first paper, "Trust the Right Teacher," proposes quality-aware self…
-
New Trajectory-Refined Distillation improves LLM training
Researchers have introduced Trajectory-Refined Distillation (TRD), a new method to improve the post-training process for large language models. TRD addresses a problem called "prefix failure" in on-policy distillation, …
-
New distillation method enhances AI safety without sacrificing reasoning
Researchers have developed a new method called Constitutional On-Policy Safe Distillation (COPSD) to improve the safety and helpfulness of AI models. Existing on-policy self-distillation techniques can lead to a collaps…