Multilingual E5
PulseAugur coverage of Multilingual E5 — every cluster mentioning Multilingual E5 across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New Nepali passport QA dataset boosts retrieval performance
Researchers have developed a new question-answering dataset specifically for Nepali passport-related services, addressing the scarcity of resources for low-resource languages. The dataset was used to fine-tune transform…
-
Multilingual embeddings show promise for translation error detection
A new study published on arXiv evaluates the effectiveness of multilingual sentence embeddings in detecting translation errors between English and Greek. Researchers developed a dataset of 1,850 examples, categorizing e…
-
Sinhala-Tamil CLIR research favors embedding models over translation
A new research paper evaluates cross-lingual information retrieval (CLIR) methods for accessing English government information using Sinhala and Tamil queries. The study compared query translation techniques, including …
-
Bekko Embedding achieves competitive multilingual retrieval with ultra-compact models
Researchers have developed Bekko Embedding, a new family of parameter-efficient multilingual retrieval models. The smallest version, bekko-embedding-v1-a8m, with under 8 million active parameters, achieves a score of 56…
-
UOL@IDEM details L1-aware vocabulary difficulty prediction for BEA 2026 task
Researchers from UOL@IDEM have detailed their submission for the BEA 2026 shared task on L1-aware vocabulary difficulty prediction. Their approach models the task as a regression problem, training separate systems for S…
-
New Slovak Text Embedding Benchmark and Models Released
Researchers have introduced SkMTEB, a new benchmark designed to evaluate text embedding models specifically for the Slovak language. This benchmark includes 31 datasets across 7 task types, significantly expanding cover…