PulseAugur
EN
LIVE 07:16:18
ENTITY On-policy self-distillation

On-policy self-distillation

PulseAugur coverage of On-policy self-distillation — every cluster mentioning On-policy self-distillation across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
7
17 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
7
17 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

5 day(s) with sentiment data

RECENT · PAGE 1/1 · 19 TOTAL
  1. TOOL · CL_252081 ·

    New SCOPE-OPSD method improves AI model self-distillation

    Researchers have developed a new method called SCOPE-OPSD, which enhances on-policy self-distillation (OPSD) by incorporating Fisher-conditioned privileged subspaces. This technique aims to improve the transfer of super…

  2. RESEARCH · CL_227065 ·

    VISTA method enhances AI reasoning via teacher-student adaptation · 2 sources tracked

    Researchers have developed VISTA, a novel method for on-policy self-distillation (OPSD) that enhances AI model reasoning. Unlike standard OPSD, VISTA adapts the teacher model based on verified student rollouts, particul…

  3. TOOL · CL_227840 ·

    New TTPO method enhances LLM math reasoning without labels

    Researchers have developed Test-Time Policy Optimization (TTPO), a novel method for improving large language models' mathematical reasoning capabilities without relying on ground-truth labels. TTPO addresses the fragili…

  4. RESEARCH · CL_223238 ·

    New TTPO method enhances LLM reasoning without labels

    Researchers have developed Test-Time Policy Optimization (TTPO), a novel method for improving large language models' mathematical reasoning capabilities without relying on ground-truth labels. TTPO addresses the fragili…

  5. RESEARCH · CL_241409 ·

    New research refines on-policy distillation for AI reasoning models

    Researchers are exploring on-policy distillation (OPD) for training reasoning models, a technique that uses a teacher model to provide per-token supervision. However, the effectiveness and optimal configuration of OPD r…

  6. TOOL · CL_212193 ·

    ArmorOCR framework enhances adversarial OCR perception with new AdvSpot benchmark

    Researchers have introduced ArmorOCR, a novel two-stage training framework designed to enhance the robustness of optical character recognition (OCR) against adversarial attacks. This framework addresses the limitations …

  7. RESEARCH · CL_193314 ·

    New framework unifies on-policy self-distillation for LLM reasoning · 3 sources tracked

    Researchers have developed a unified framework for on-policy self-distillation (OPSD) to enhance LLM reasoning by integrating privileged information into model parameters. This new framework, Unified On-Policy Self-Dist…

  8. TOOL · CL_183142 ·

    New AI training method uses rubrics as privileged information for open-ended generation

    Researchers have developed a new method called On-policy self-distillation (OPSD) that utilizes rubrics as privileged information (PI) for open-ended text generation. This approach enhances the training signal for model…

  9. TOOL · CL_179300 ·

    New method CSCR improves LLM long-context reasoning by reallocating token credit

    Researchers have developed a new method called Counterfactual Sensitivity Credit Reallocation (CSCR) to improve the reasoning capabilities of large language models, particularly in tasks requiring long-context reasoning…

  10. RESEARCH · CL_160792 ·

    Visual Contrastive Self-Distillation Improves Qwen VL Models

    Researchers have developed Visual Contrastive Self-Distillation (VCSD), a novel method for improving Vision-Language Models (VLMs) without requiring external teachers or privileged information. VCSD works by comparing a…

  11. TOOL · CL_156456 ·

    New distillation method trains LLMs efficiently with soft prompts

    Researchers have developed a new method called Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context ("method") to train large language models. This technique uses a teacher model that differs from the st…

  12. TOOL · CL_154090 ·

    New research identifies decoding collapse in AI agent self-distillation

    Researchers have identified a failure mode in feedback-augmented self-distillation for retrieval-interleaved search agents, termed decoding collapse. This occurs when models generate diverse-looking but input-agnostic r…

  13. RESEARCH · CL_156464 ·

    New H$^2$SD framework boosts LLM reasoning via hybrid self-distillation

    Researchers have developed H$^2$SD, a novel hybrid hindsight self-distillation framework designed to enhance the reasoning abilities of large language models. This method addresses limitations in existing reinforcement …

  14. RESEARCH · CL_167450 ·

    New on-policy distillation methods enhance LLM reasoning and efficiency · 10 sources tracked

    Multiple research papers explore advancements in on-policy distillation (OPD) techniques for language models, aiming to improve reasoning capabilities and training efficiency. Several methods, including SimpleOPD, S$^2$…

  15. RESEARCH · CL_141172 ·

    New research tackles pathologies in On-Policy Distillation for LLMs

    Researchers have identified and proposed solutions for two key pathologies in On-Policy Distillation (OPD), a technique used in large language model post-training. The first pathology, Student-Teacher Mismatch, occurs w…

  16. RESEARCH · CL_117125 ·

    New research challenges on-policy self-distillation for LLMs, proposing refined methods · 10 sources tracked

    Recent research papers explore the limitations and potential improvements of on-policy self-distillation (OPSD) for training large language models (LLMs). Studies indicate that standard OPSD can lead to rote memorizatio…

  17. RESEARCH · CL_90827 ·

    New methods enhance VLM accuracy for GUI grounding tasks · 2 papers

    Two new research papers introduce novel methods for improving the accuracy and reliability of vision-language models (VLMs) in GUI grounding tasks. The first paper, "Trust the Right Teacher," proposes quality-aware self…

  18. RESEARCH · CL_79119 ·

    New Trajectory-Refined Distillation improves LLM training

    Researchers have introduced Trajectory-Refined Distillation (TRD), a new method to improve the post-training process for large language models. TRD addresses a problem called "prefix failure" in on-policy distillation, …

  19. TOOL · CL_68337 ·

    New distillation method enhances AI safety without sacrificing reasoning

    Researchers have developed a new method called Constitutional On-Policy Safe Distillation (COPSD) to improve the safety and helpfulness of AI models. Existing on-policy self-distillation techniques can lead to a collaps…