PulseAugur
中
实时 06:23:56

新方法改进策略内蒸馏以增强人工智能能力 · 跟踪 7 个来源

研究人员正在探索策略内蒸馏(OPD)的高级技术,通过结合多个“教师”模型来增强语言模型的能力。LEGO-OPD 和 SAKI 等新方法侧重于分解和路由教师信号,以在不降低性能的情况下改善视觉基础和推理。SCOUT 和 MAESTRO 等其他方法解决了教师-学生前缀不匹配的挑战,并优化了教师干预策略,以防止性能下降并确保跨不同任务的平衡能力集成。 AI

影响 策略内蒸馏的这些进展可能带来更强大、更专业的人工智能模型,从而提高复杂推理和多模态任务的性能。

排序理由 多篇研究论文介绍了策略内蒸馏的新技术。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 25 个来源。 我们如何撰写摘要 →

新方法改进策略内蒸馏以增强人工智能能力 · 跟踪 7 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了策略内蒸馏的新技术。
Source corroboration
25 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
13 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+10 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [25]

  1. arXiv cs.LG TIER_1 English(EN) · Hejian Sang, Zhengze Zhou, Shayan Mohajer Hamidi, Xiaomin Li, Rohit Jain, Alborz Geramifard ·

    教师各自所学:通过教师相对偏移进行多教师策略内蒸馏

    arXiv:2610.10460v1 Announce Type: new Abstract: Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillat…

  2. arXiv cs.LG TIER_1 English(EN) · Randy Ardywibowo, Arnav Dalal, Jiantao Jiao ·

    善于因材施教的自学教师:联合策略学习与教学

    arXiv:2610.10447v1 Announce Type: new Abstract: Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an at…

  3. arXiv cs.CL TIER_1 English(EN) · Yixuan Tang, Yi Yang ·

    On-Policy Distillation Teaches New Skills but Not New Knowledge

    arXiv:2610.09639v1 Announce Type: new Abstract: On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled…

  4. arXiv cs.AI TIER_1 English(EN) · Xiaofei Yin, Tong Chu, Jiyuan Fu, Jun Lan, Shuheng Zhou, Huijia Zhu ·

    面向能力保持的慢快多教师在线策略蒸馏

    arXiv:2610.02324v1 Announce Type: cross Abstract: Foundation multimodal large language models are designed to support a broad spectrum of capabilities across diverse domains. Multi-teacher on-policy distillation (MOPD) provides an effective framework for consolidating domain-spec…

  5. arXiv cs.AI TIER_1 English(EN) · Xiang Chen, Futao Su, Kong Wang, Jiayi Chen, TanLin Li ·

    将教师Token用在关键之处:基于成功的On-Policy蒸馏

    arXiv:2610.02678v1 Announce Type: new Abstract: On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher, but providing such supervision for every rollout requires substantial teacher computation. We introduce Success-Refe…

  6. arXiv cs.LG TIER_1 English(EN) · Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li ·

    Latent-MOPD:潜在多教师策略内蒸馏

    arXiv:2610.02381v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first …

  7. arXiv cs.LG TIER_1 English(EN) · Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang, Zhanyang Jin, Yihang Sun, Jiaxuan You ·

    从梯度到能力:理解多教师策略内蒸馏

    arXiv:2610.02179v1 Announce Type: new Abstract: Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teach…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    DiffGate:难度门控教师指导用于在线策略蒸馏

    On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely…

  9. arXiv cs.AI TIER_1 English(EN) · Jaeyun Shin, Hangeol Chang, Jong Chul Ye ·

    LEGO-OPD:多模态 on-policy 蒸馏的因子化教师组合

    arXiv:2610.00333v1 Announce Type: cross Abstract: Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary …

  10. arXiv cs.AI TIER_1 English(EN) · Langlin Huang, Hao Liu, Mononito Goswami, Xinyu Li, Prithwith Jana, Nikos Kanakaris, Patrick Bl\"obaum, Purak Jain ·

    On-Policy Distillation中的Off-Policy Teacher

    arXiv:2609.38360v1 Announce Type: cross Abstract: On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundament…

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    从梯度到能力:理解多教师策略内蒸馏

    Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    Latent-MOPD:潜在多教师策略内蒸馏

    On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method fo…

  13. arXiv cs.CL TIER_1 English(EN) · Yuhao Wang, Ruiyang Ren, Yinan Zhang, Ruiqing Zhang, Jing Liu, Chunyan Miao ·

    从不和谐到协同:教师干预在策略内蒸馏中的应用

    arXiv:2609.37510v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student lear…

  14. arXiv cs.AI TIER_1 English(EN) · Miteto Wei, Xiaohan Wang, Zehao Chen, Jiajun Chai, Sichao Liu, Li Wang, Haoyuan Xu, Zhaoyu Hu, Wei Lin, Guojun Yin ·

    SAKI: 用于在线策略蒸馏的最大耦合路由教师监督

    arXiv:2609.36601v1 Announce Type: new Abstract: On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce S…

  15. Hugging Face Daily Papers TIER_1 English(EN) ·

    LEGO-OPD:多模态 on-policy 蒸馏的因子化教师组合

    Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full …

  16. arXiv cs.AI TIER_1 English(EN) · Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma, Jihuai Zhu, Lelun Wu, Lyumanshan Ye, Pengfei Liu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu ·

    MOPD-Router:重新思考多教师在线策略蒸馏中的教师路由

    arXiv:2609.30837v2 Announce Type: cross Abstract: Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on …

  17. Hugging Face Daily Papers TIER_1 English(EN) ·

    On-Policy Distillation中的Off-Policy Teacher

    On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories ar…

  18. Hugging Face Daily Papers TIER_1 English(EN) ·

    SAKI: 用于在线策略蒸馏的最大耦合路由教师监督

    On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained …

  19. Hugging Face Daily Papers TIER_1 English(EN) ·

    PMOPD:多教师策略内蒸馏中的任务排序、循环和参数更新子空间保护

    Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillat…

  20. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越教师分配:域归一化多教师 on-policy 蒸馏

    Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letti…

  21. Hugging Face Daily Papers TIER_1 English(EN) ·

    PMOPD:多教师策略内蒸馏中的任务排序、循环和参数更新子空间保护

    Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillat…

  22. Hugging Face Daily Papers TIER_1 English(EN) ·

    DuoOPD:从联合师生结果中学习以进行多任务在线策略蒸馏

    On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and…

  23. Hugging Face Daily Papers TIER_1 English(EN) ·

    MOPD-Router:重新思考多教师在线策略蒸馏中的教师路由

    Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabel…

  24. arXiv cs.CV TIER_1 English(EN) · Taojie Zhu, Jing Jin, Yuan Xia, Chenyang Ding, Qunshan He, Wanke Xia, Tao Sun, Yan Chen, Jian Wang, Jinjie Gu, Tao Feng ·

    UP-MOPD:多教师在线策略蒸馏中的更新投影

    arXiv:2610.08398v1 Announce Type: new Abstract: On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration. Gradient corrections directly constrain parameter updates under plai…

  25. arXiv cs.CV TIER_1 English(EN) · Siyuan Liu, Kanghui Tian, Yue Duan, Yutao He, Shangdong Yang, Jian Zhang, Yinghuan Shi ·

    更好的教师监督是否足够?在多模态按策略蒸馏中解锁学生端的学习

    arXiv:2609.39120v1 Announce Type: new Abstract: On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods primarily focus on enhancing this teacher-side guidance (e.g., by enriching teache…