Researchers have introduced XHotpotQA, a new benchmark designed to evaluate cross-lingual knowledge composition in multi-hop question answering systems. Unlike previous benchmarks that translate entire examples, XHotpotQA explicitly assigns languages to different components of the question-answering process, including the question, bridge evidence, and answer-bearing evidence. The benchmark includes 15,661 training and 7,405 validation instances, with detailed sentence-level support supervision and supplied distractors. Initial evaluations show that mismatches across language and script interfaces significantly degrade performance, highlighting the challenges for AI systems that need to integrate information from diverse linguistic sources. AI
IMPACT This benchmark will help researchers develop AI systems capable of integrating information across different languages, crucial for global knowledge access.
RANK_REASON The item describes a new academic benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
- XHotpotQA
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →