PulseAugur
实时 10:48:28
English(EN) Behavior Leverage Imbalance in Multi-Teacher On-Policy Distillation

新的Soft Clamp方法解决了代理语言模型的过度调用问题

研究人员在代理语言模型的多教师同策略蒸馏中发现了一个“行为杠杆失衡”问题。这种失衡会导致模型过度调用工具,即使在直接回答更合适的情况下也是如此,而这个问题无法通过聚合损失指标来检测。为了解决这个问题,提出了一种名为Soft Clamp的新方法,该方法对每令牌发散进行校准以减轻极端信号。实验表明,Soft Clamp显著减少了过度调用,并提高了APIGen-MT等基准测试的决策准确性。 AI

影响 这项研究可能有助于开发更可靠的代理语言模型,使其能更好地理解何时使用工具以及何时提供直接答案。

排序理由 该集群包含一篇详细介绍新AI模型训练方法的学术论文。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新的Soft Clamp方法解决了代理语言模型的过度调用问题

报道来源 [3]

  1. arXiv cs.CL TIER_1 English(EN) · Jiabin Shen, Guang Chen, Chengjun Mao ·

    多教师同策略蒸馏中的行为杠杆失衡

    arXiv:2607.07050v1 Announce Type: new Abstract: Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly. This makes multi-teacher on-policy distillation a natural training strategy: one teacher can specialize in tool call…

  2. arXiv cs.CL TIER_1 English(EN) · Chengjun Mao ·

    多教师On-Policy蒸馏中的行为杠杆失衡

    Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly. This makes multi-teacher on-policy distillation a natural training strategy: one teacher can specialize in tool calls, another in direct responses, and the student …

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    多教师同策略蒸馏中的行为杠杆失衡

    Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly. This makes multi-teacher on-policy distillation a natural training strategy: one teacher can specialize in tool calls, another in direct responses, and the student …