PulseAugur
实时 21:14:33
English(EN) Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation

新的TAPO方法通过显式纠错增强LLM自蒸馏 · 跟踪4个来源

研究人员推出了一种新方法,称为轨迹增强策略优化(TAPO),用于大型语言模型的自蒸馏。与隐式对齐分布的传统方法不同,TAPO显式地构建了纠正性轨迹。这些轨迹保留了错误推理直到失败点,然后纳入自然语言诊断和纠正后的推理。在AIME 2024、AIME 2025和HMMT 2025上的实验表明,与GRPO相比,TAPO提高了初始推理和纠错的有效性。 AI

影响 通过在训练过程中提供更有针对性的纠错来增强LLM的推理能力。

排序理由 该集群描述了一篇关于通过自蒸馏改进大型语言模型推理的新颖方法的新研究论文。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新的TAPO方法通过显式纠错增强LLM自蒸馏 · 跟踪4个来源

报道来源 [3]

  1. arXiv cs.LG TIER_1 English(EN) · Zhilin Huang, Hang Gao, Ziqiang Dong, Yuan Chen, Yifeng Luo, Chujun Qin, Jingyi Wang, Yang Yang, Guanjun Jiang ·

    从错误中学习:为自蒸馏构建可学习的微反射轨迹

    arXiv:2606.18844v1 Announce Type: new Abstract: Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes KL divergence toward a privileged target distributio…

  2. arXiv cs.LG TIER_1 English(EN) · Guanjun Jiang ·

    从错误中学习:为自蒸馏构建可学习的微反射轨迹

    Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes KL divergence toward a privileged target distribution. However, because this supervision is generate…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    从错误中学习:构建可学习的微反思轨迹以实现自蒸馏

    Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes KL divergence toward a privileged target distribution. However, because this supervision is generate…