A new arXiv paper titled "Actions Speak Louder than Words" investigates how tool-using AI agents perform tasks across different languages. The research highlights that simply comparing final answers is insufficient, as the sequence of actions taken by the agent is crucial for understanding performance, cost, and failure modes. The study analyzed over 2.38 million rollouts across 8 models and 41 languages, revealing significant cross-lingual divergence in action policies that is structural rather than random noise. The findings suggest that many frontier models tend to route non-English tasks through English, a behavior that persists even when instructed otherwise. AI
IMPACT Reveals critical limitations in cross-lingual reasoning for tool-using AI agents, impacting global deployment and evaluation.
RANK_REASON Research paper published on arXiv detailing methodology and findings. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →