Belebele
PulseAugur coverage of Belebele — every cluster mentioning Belebele across labs, papers, and developer communities, ranked by signal.
-
LLM language support claims differ from official commitments
Large language models often claim to support over a hundred languages due to their extensive pretraining data, but this fluency does not equate to official support or verified benchmark performance. Vendors like Meta (L…
-
AI benchmarks cover less than 3% of world languages, study finds
Current AI language model benchmarks significantly underrepresent the world's linguistic diversity, with the broadest benchmarks covering only about 2.9% of the roughly 7,000 living languages. Even the most comprehensiv…
-
Indian languages face 8x "tokenizer tax" in LLMs due to English-centric training
A new research paper highlights a significant disadvantage faced by Indian languages when processed by large language models due to subword tokenization. These tokenizers, primarily trained on English data, result in an…
-
LangMAP tokenization adapts multilingual models without vocabulary changes
Researchers have developed LangMAP, a novel language-adaptive tokenization approach that extends the UnigramLM algorithm for multilingual settings. This method allows for language-specific tokenization from a single sha…
-
Multilingual Code-Switching Boosts LLM Performance Across Four Languages
Researchers have explored the impact of multilingual code-switching data (CSD) on large language models (LLMs) across four languages: English, Japanese, Korean, and Chinese. Their experiments demonstrated that incorpora…
-
New research tackles multilingual adaptation in Mixture-of-Experts models
Two new research papers explore the adaptation of Mixture-of-Experts (MoE) models for multilingual tasks. One paper analyzes how language specialization emerges in MoE models during continual pre-training, finding that …