Researchers have developed a BETA-labeling framework to construct multilingual datasets for low-resource information retrieval (IR). This framework utilizes multiple large language models (LLMs) with checks for consistency and majority agreement, followed by human evaluation to ensure label quality. The study also investigated the feasibility of reusing IR datasets from other low-resource languages through machine translation, revealing significant variations and semantic preservation issues that impact the reliability of cross-lingual dataset reuse. The findings highlight both the potential and limitations of LLM-assisted dataset creation for low-resource IR, offering guidance for building more dependable benchmarks. AI
IMPACT Provides methods to improve AI model performance in low-resource languages, potentially expanding AI accessibility.
RANK_REASON Academic paper detailing a new methodology for dataset construction. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →