XLM-RoBERTa
PulseAugur coverage of XLM-RoBERTa — every cluster mentioning XLM-RoBERTa across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
Domain-specific pretraining boosts Transformer performance on Arabic-English code-switching
A new study published on arXiv explores the impact of domain-specific pretraining on Transformer models for analyzing Arabic-English code-switching. The research evaluated MARBERT and XLM-RoBERTa, with BERT as a baselin…
-
AI framework incorporates annotator psychology for sexism detection
Researchers from VANGUARD have developed a multimodal framework for detecting sexism online, incorporating annotator psychology and demographics into the detection process. Their approach fuses five input modalities usi…
-
New IndicTriMix method improves language identification in code-mixed text
Researchers have developed a new method called IndicTriMix for identifying languages within code-mixed text, which is common in social media. This approach treats language identification as a sequence labeling problem a…
-
Transformers Mimic Traditional Models in Multilingual Readability Assessment
Researchers have analyzed how Transformer-based models and traditional feature-based models approach multilingual readability assessment. They found that while Transformers achieve high accuracy, their internal feature …
-
Fine-tuning LLMs: Four crucial steps before you start
The article advises against immediately fine-tuning large language models like GPT-3, Bert, T5, Roberta, and XLM-RoBERTa. It suggests performing four crucial steps before proceeding with fine-tuning to ensure better and…
-
Hugging Face unveils efficient multimodal encoder NeoMME, study favors encoders for Indic NER
Hugging Face has introduced NeoMME, a new family of multilingual multimodal encoders designed for efficiency. Unlike many generative models, NeoMME uses a single bidirectional Transformer to process both text and image …
-
New TabuLM model enhances low-resource language processing with tabular data
Researchers have developed TabuLM, a novel language model specifically pre-trained on Kinyarwanda tabular data to address the scarcity of resources for low-resource languages. This model extends KinyaBERT-large by incor…
-
Prompt compression struggles with non-English languages, study finds
A new study published on arXiv titled "Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors" investigates the effectiveness of prompt compression techniques across different languages. …
-
TabuLM: First Language Model Pre-trained on Kinyarwanda Tabular Data
Researchers have developed TabuLM, a novel language model specifically pre-trained on tabular data for Kinyarwanda, a low-resource Bantu language spoken in Rwanda. This model enhances KinyaBERT-large with new embeddings…
-
Hugging Face Model Selection Framework Launched Amidst 2M Model Milestone
This article provides a framework for selecting and fine-tuning models from Hugging Face, a platform that now hosts over two million models. It guides users through the decision-making process, offering insights beyond …
-
Sparse few-shot language model for Bengali achieves 90% sparsity
Researchers have developed BnBERT-iPET, a novel approach to sparse few-shot language modeling specifically for Bengali. This method utilizes lottery ticket pruning to achieve 90% sparsity, significantly reducing computa…
-
OpenAI's Privacy Filter shows mixed results in PII detection study
A new research paper introduces the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5 billion parameter PII detector. The study found that OPF performs well on structured PII types like emails and phon…
-
Kazakh-Russian code-switching identification bottleneck is annotation, not model
A new paper on arXiv explores the identification of code-switching between Kazakh and Russian languages, finding that the annotation boundary is more critical than the model used. Researchers developed a gold LID (Langu…
-
BioSentinel uses XLM-RoBERTa for sexism intent classification in memes
The BioSentinel team developed a text-centric approach for classifying sexist intent in memes for the EXIST 2026 competition. Their system, built on XLM-RoBERTa-base, utilized a combined loss function incorporating KL d…
-
Indian languages face 8x "tokenizer tax" in LLMs due to English-centric training
A new research paper highlights a significant disadvantage faced by Indian languages when processed by large language models due to subword tokenization. These tokenizers, primarily trained on English data, result in an…
-
Prompt compression fails non-English languages, new paper finds
A new paper titled "Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors" reveals that prompt compression techniques, designed to reduce LLM inference costs by removing low-information …
-
New dataset and multimodal model tackle Bengali YouTube clickbait
Researchers have introduced BanClickThumb, a new multimodal dataset designed to detect clickbait in Bengali YouTube videos. The dataset comprises 7,147 thumbnail-title pairs and was used to benchmark various detection m…
-
New study compares AI strategies for multilingual polarization detection
Researchers have conducted a comparative study on multilingual polarization detection across 22 languages for SemEval-2026 Task 9. The study evaluated generalist models, language-specific specialists, and ensemble strat…
-
Urdu fake news detection hindered by dataset length confound
Researchers have conducted a study on Urdu fake news detection, highlighting significant challenges in cross-dataset generalization. Using the XLM-RoBERTa model and two distinct Urdu datasets, the study found that while…
-
New VEXMLM model boosts performance for African languages
Researchers have developed VEXMLM, a new language model designed to improve performance on low-resource African languages, specifically Amharic and Tigrinya, which use the Ge'ez script. This model addresses issues of hi…