PulseAugur
实时 09:10:40
English(EN) Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

研究发现:AI代理的跨语言行动策略存在显著差异

一篇新的arXiv论文,题为“行动胜于雄辩”,研究了使用工具的AI代理如何在不同语言中执行任务。研究强调,仅仅比较最终答案是不够的,因为代理所采取的行动顺序对于理解性能、成本和失败模式至关重要。该研究分析了8个模型和41种语言的超过238万次运行,揭示了行动策略中存在显著的跨语言差异,这种差异是结构性的,而非随机噪声。研究结果表明,许多前沿模型倾向于将非英语任务通过英语路由,即使在被指示采取其他方式时,这种行为仍然存在。 AI

影响 揭示了使用工具的AI代理在跨语言推理方面的关键局限性,影响其全球部署和评估。

排序理由 发表在arXiv上的研究论文,详细介绍了方法和研究结果。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

研究发现:AI代理的跨语言行动策略存在显著差异

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram ·

    行动胜于雄辩:衡量多语言策略在工具使用代理中的保留情况

    arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: t…