A new research paper evaluates cross-lingual information retrieval (CLIR) methods for accessing English government information using Sinhala and Tamil queries. The study compared query translation techniques, including Google Translate, NLLB-200, and mBART50, against cross-lingual embedding models like LaBSE, Multilingual E5, and BGE-M3. Experiments on a benchmark dataset from Sri Lanka's Government Information Center showed that while both approaches improved retrieval accuracy over monolingual methods, the BGE-M3 embedding model achieved the highest performance, demonstrating its effectiveness for low-resource government domains. AI
IMPACT Demonstrates the superiority of embedding models over translation for low-resource cross-lingual retrieval, potentially improving access to information in underserved domains.
RANK_REASON Academic paper on cross-lingual information retrieval methods. [lever_c_demoted from research: ic=1 ai=1.0]
Read on arXiv cs.IR (Information Retrieval) →
- BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- English
- Google Translate
- Government Information Center
- LaBSE
- mBART50
- Multilingual E5
- NLLB-200
- Sinhala
- Sri Lanka
- Tamil
- Tiroshan Madushanka
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →