一篇新研究论文评估了使用僧伽罗语和泰米尔语查询访问英文政务信息的跨语言信息检索(CLIR)方法。该研究比较了查询翻译技术(包括 Google Translate、NLLB-200 和 mBART50)与跨语言嵌入模型(如 LaBSE、Multilingual E5 和 BGE-M3)的性能。在斯里兰卡政府信息中心的基准数据集上进行的实验表明,虽然两种方法都比单语方法提高了检索准确性,但 BGE-M3 嵌入模型取得了最高的性能,证明了其在低资源政务领域的有效性。 AI
影响 证明了嵌入模型在低资源跨语言检索方面优于翻译模型,可能改善对服务不足领域信息的访问。
排序理由 关于跨语言信息检索方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
在 arXiv cs.IR (Information Retrieval) 阅读 →
- BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- English
- Google Translate
- Government Information Center
- LaBSE
- mBART50
- Multilingual E5
- NLLB-200
- Sinhala
- Sri Lanka
- Tamil
- Tiroshan Madushanka
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →