LaBSE
PulseAugur coverage of LaBSE — every cluster mentioning LaBSE across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
QuanLing framework quantifies language distance in Western Romance languages
Researchers have extended the QuanLing framework, which uses pretrained language models to quantify language distance, to Western Romance languages. This framework combines sentence embedding distances, tokenization fra…
-
Multilingual embeddings show promise for translation error detection
A new study published on arXiv evaluates the effectiveness of multilingual sentence embeddings in detecting translation errors between English and Greek. Researchers developed a dataset of 1,850 examples, categorizing e…
-
KinyaEmbed model enhances Kinyarwanda language processing with novel training
Researchers have developed KinyaEmbed, a new sentence embedding model specifically designed for the Kinyarwanda language. This model addresses the poor performance of existing multilingual models on Kinyarwanda due to i…
-
Wikipedia Language Editions Show Divergence in Religion Framing
A new study analyzed framing differences across Wikipedia's language editions for matched concepts, finding that scientific articles exhibit closer alignment than calibration concepts. The research utilized embedding di…
-
Sinhala-Tamil CLIR research favors embedding models over translation
A new research paper evaluates cross-lingual information retrieval (CLIR) methods for accessing English government information using Sinhala and Tamil queries. The study compared query translation techniques, including …
-
UOL@IDEM details L1-aware vocabulary difficulty prediction for BEA 2026 task
Researchers from UOL@IDEM have detailed their submission for the BEA 2026 shared task on L1-aware vocabulary difficulty prediction. Their approach models the task as a regression problem, training separate systems for S…
-
New method uses LaBSE embeddings for cross-lingual polarization detection
Researchers have developed a novel approach to detect online polarization across multiple languages and cultures, addressing the challenge of limited data in low-resource languages. Their method utilizes LaBSE embedding…
-
OmniSONAR models process thousands of languages in text and speech
Researchers have introduced OmniSONAR, a novel family of sentence embedding models capable of processing thousands of languages across text and speech. This system achieves state-of-the-art performance by using progress…
-
Open Diachronic Greek Treebank Released with Indo-European Parallels
Researchers have developed AthDGC, a comprehensive, open-source dataset and workflow for dependency parsing of the Greek language across eight historical periods. This project, built upon the PROIEL Treebank Family sche…
-
Google Embeddings 2 leads retrieval benchmarks but lags in speed
A new paper benchmarks Google Embeddings 2 (GE2) against several open-source models for multilingual dense retrieval and RAG systems. GE2 achieved top performance across multiple tasks, including BEIR and an Italian RAG…
-
Machine translation preserves moral semantics across languages
Researchers have demonstrated that machine translation, particularly using LLMs, can effectively preserve subtle moral cues across languages. A study using approximately 50,000 morally-annotated social media posts from …