MMTEB: Massive Multilingual Text Embedding Benchmark
PulseAugur coverage of MMTEB: Massive Multilingual Text Embedding Benchmark — every cluster mentioning MMTEB: Massive Multilingual Text Embedding Benchmark across labs, papers, and developer communities, ranked by signal.
-
New benchmarks evaluate Portuguese text embedding models, revealing performance gaps
Two new benchmarks, MTEB-PT and MTEB-PT (Brazilian Portuguese), have been released to evaluate text embedding models specifically for the Portuguese language. These benchmarks address the underrepresentation of Portugue…
-
New BITEMBED framework drastically cuts LLM embedding costs
Researchers have developed BITEMBED, a novel framework designed to create efficient text embeddings for large language models. This approach converts LLM backbones into low-bit encoders using ternary weights and quantiz…
-
HAKARI-Bench offers lightweight evaluation for retrieval models · 2 sources tracked
Researchers have introduced HAKARI-Bench, a lightweight benchmark designed to streamline the evaluation of retrieval architectures and efficiency settings for retrieval-augmented generation and semantic search. This new…
-
New method corrects mean bias in text embeddings
Researchers have identified a consistent bias in current text embedding models, where each embedding can be decomposed into a sentence-specific component and a near-identical mean component across all sentences. They pr…