A new benchmark called SWORD has been developed to evaluate Large Language Models' (LLMs) ability to reject factual errors across different languages. SWORD uses distortions derived from Wikidata to create factually incorrect statements, revealing that models perform better on semantically plausible errors than random ones. The benchmark also highlights significant performance disparities in LLMs when processing East Asian languages compared to others, with accuracy drops of up to 28 percentage points. AI
IMPACT Highlights critical cross-lingual weaknesses in LLM factual reasoning, suggesting current benchmarks may obscure true understanding.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →