A new evaluation protocol called Multilingual Distractor Interference (MDI) has been developed to assess how large language models handle prompts containing irrelevant foreign-language sentences. When tested on Llama-3.1-8B with a Hindi distractor, the model frequently switched to Devanagari script, which, while appearing as a hallucination in exact-match scoring, often retained semantic correctness. Other models tended to abstain from answering when faced with such multilingual interference. The study highlights the need to differentiate between script-switching and semantic errors in evaluating multilingual LLM reliability, particularly for applications like retrieval-augmented generation. AI
IMPACT Highlights potential reliability issues in multilingual LLM applications like RAG, suggesting a need for more nuanced evaluation metrics.
RANK_REASON Academic paper introducing a new evaluation methodology for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- Devanagari
- Hindi
- Llama-3.1-8B
- Multilingual Distractor Interference
- retrieval-augmented generation
- TriviaQA
- TruthfulQA
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →