PulseAugur
实时 15:10:19
English(EN) Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

新研究解决了大型语言模型在线策略蒸馏中的病理问题

研究人员已识别出在线策略蒸馏(OPD)中的两个关键病理问题,并提出了解决方案。OPD是大型语言模型(LLM)后训练中使用的一种技术。第一个病理问题是学生-教师模型不匹配,当教师模型和学生模型之间存在显著差距时,会导致指导失准。第二个是长度利用,当模型学会操纵响应长度以获得更高奖励时出现。为解决这些问题,引入了优势裁剪、对数尺度压缩和自适应双视角OPD(AD-OPSD)等新方法来调节蒸馏信号并保留模型的原生推理能力,在基准测试中显示出更高的准确性。 AI

影响 这些发现可能导致更稳定、更有效的LLM训练,提高复杂推理任务的性能。

排序理由 该集群包含两篇arXiv论文,详细介绍了LLM训练方法的研究并提出了新技术。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

新研究解决了大型语言模型在线策略蒸馏中的病理问题

报道来源 [6]

  1. arXiv cs.LG TIER_1 English(EN) · Rui Wang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Wenhao Yu, Kam-Fai Wong ·

    揭秘 On-Policy Distillation:角色、病理和法规

    arXiv:2607.13399v1 Announce Type: cross Abstract: On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clari…

  2. arXiv cs.CL TIER_1 English(EN) · Kam-Fai Wong ·

    揭秘 On-Policy Distillation:作用、病理和监管

    On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it …

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    揭秘On-Policy Distillation:角色、病理和法规

    On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it …

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    揭秘On-Policy Distillation:角色、病理与监管

    On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it …

  5. arXiv cs.CL TIER_1 English(EN) · Keqin Peng, Chen Li, Yuanxin Ouyang, Yancheng Yuan, Liang Ding ·

    诊断和缓解在线策略自蒸馏中的思维崩溃

    arXiv:2607.10805v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxically degrades downstream performance. In this paper, we…

  6. arXiv cs.CL TIER_1 English(EN) · Liang Ding ·

    诊断和缓解在线策略自蒸馏中的思维崩溃

    On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxically degrades downstream performance. In this paper, we systematically investigate this pathology and i…