PulseAugur
EN
LIVE 10:41:11

Study reveals tool-using agents often pivot to English, impacting cross-lingual performance

A new study, "Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents," evaluated how tool-using agents perform when given tasks in different languages. The research found that while previous multilingual evaluations focused on final answers, the actual actions taken by the agents are crucial for understanding cost, latency, and failure modes. The study analyzed 2.38 million rollouts across 8 models, 6 benchmarks, and 41 languages, identifying and correcting five confounds that previously obscured true performance. AI

IMPACT Reveals that tool-using agents often default to English even when prompted in other languages, highlighting a critical area for improvement in multilingual AI capabilities.

RANK_REASON Research paper analyzing model behavior. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Study reveals tool-using agents often pivot to English, impacting cross-lingual performance

COVERAGE [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

    When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how the system …