A new research paper explores the challenge of multilingual tool use in large language models (LLMs), specifically addressing the issue of "Argument Language Mismatch" (ALM). This problem occurs when an LLM correctly identifies a tool but generates arguments in an inconsistent language, rendering the output invalid. The study found that supervised fine-tuning (SFT) is a highly effective method for mitigating ALM, significantly improving argument language consistency and overall function call accuracy. While reinforcement learning (RL) approaches, like Group Relative Policy Optimization (GRPO), offer incremental gains in language consistency and general reasoning, SFT provides a strong baseline and often comparable performance. AI
IMPACT Addresses a key limitation in LLM tool use, potentially improving reliability for multilingual applications.
RANK_REASON Research paper published on arXiv detailing a novel approach to improving LLM multilingual tool use. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- Argument Language Mismatch
- arXiv
- CatalyzeX
- DagsHub
- Group Relative Policy Optimization
- Hugging Face
- large-language models
- reinforcement learning
- supervised fine-tuning
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →