PulseAugur
EN
LIVE 06:38:32

Sinhala-Tamil CLIR research favors embedding models over translation

A new research paper evaluates cross-lingual information retrieval (CLIR) methods for accessing English government information using Sinhala and Tamil queries. The study compared query translation techniques, including Google Translate, NLLB-200, and mBART50, against cross-lingual embedding models like LaBSE, Multilingual E5, and BGE-M3. Experiments on a benchmark dataset from Sri Lanka's Government Information Center showed that while both approaches improved retrieval accuracy over monolingual methods, the BGE-M3 embedding model achieved the highest performance, demonstrating its effectiveness for low-resource government domains. AI

IMPACT Demonstrates the superiority of embedding models over translation for low-resource cross-lingual retrieval, potentially improving access to information in underserved domains.

RANK_REASON Academic paper on cross-lingual information retrieval methods. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.IR (Information Retrieval) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Sinhala-Tamil CLIR research favors embedding models over translation

COVERAGE [1]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Tiroshan Madushanka ·

    Query Translation vs. Cross-Lingual Embeddings for Sinhala-Tamil E-Government Information Retrieval

    This paper presents a comparative evaluation of cross-lingual information retrieval (CLIR) methods for retrieving English government information using Sinhala and Tamil queries. Two CLIR paradigms are investigated: Query Translation (QT), employing Google Translate, NLLB, and mBA…