PulseAugur
中
实时 03:59:10
English(EN) Query Translation vs. Cross-Lingual Embeddings for Sinhala-Tamil E-Government Information Retrieval

僧伽罗语-泰米尔语跨语言信息检索研究:嵌入模型优于翻译模型

一篇新研究论文评估了使用僧伽罗语和泰米尔语查询访问英文政务信息的跨语言信息检索(CLIR)方法。该研究比较了查询翻译技术(包括 Google Translate、NLLB-200 和 mBART50)与跨语言嵌入模型(如 LaBSE、Multilingual E5 和 BGE-M3)的性能。在斯里兰卡政府信息中心的基准数据集上进行的实验表明,虽然两种方法都比单语方法提高了检索准确性,但 BGE-M3 嵌入模型取得了最高的性能,证明了其在低资源政务领域的有效性。 AI

影响 证明了嵌入模型在低资源跨语言检索方面优于翻译模型,可能改善对服务不足领域信息的访问。

排序理由 关于跨语言信息检索方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

僧伽罗语-泰米尔语跨语言信息检索研究:嵌入模型优于翻译模型

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
关于跨语言信息检索方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
47 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Tiroshan Madushanka ·

    查询翻译与跨语言嵌入在僧伽罗语-泰米尔语电子政务信息检索中的应用

    This paper presents a comparative evaluation of cross-lingual information retrieval (CLIR) methods for retrieving English government information using Sinhala and Tamil queries. Two CLIR paradigms are investigated: Query Translation (QT), employing Google Translate, NLLB, and mBA…