Belebele
PulseAugur coverage of Belebele — every cluster mentioning Belebele across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Indian languages face 8x "tokenizer tax" in LLMs due to English-centric training
A new research paper highlights a significant disadvantage faced by Indian languages when processed by large language models due to subword tokenization. These tokenizers, primarily trained on English data, result in an…
-
LangMAP tokenization adapts multilingual models without vocabulary changes
Researchers have developed LangMAP, a novel language-adaptive tokenization approach that extends the UnigramLM algorithm for multilingual settings. This method allows for language-specific tokenization from a single sha…
-
Multilingual Code-Switching Boosts LLM Performance Across Four Languages
Researchers have explored the impact of multilingual code-switching data (CSD) on large language models (LLMs) across four languages: English, Japanese, Korean, and Chinese. Their experiments demonstrated that incorpora…
-
New research tackles multilingual adaptation in Mixture-of-Experts models
Two new research papers explore the adaptation of Mixture-of-Experts (MoE) models for multilingual tasks. One paper analyzes how language specialization emerges in MoE models during continual pre-training, finding that …