PulseAugur
EN
LIVE 18:53:53
ENTITY XLM-RoBERTa

XLM-RoBERTa

PulseAugur coverage of XLM-RoBERTa — every cluster mentioning XLM-RoBERTa across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
10
32 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
10
30 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

9 day(s) with sentiment data

RECENT · PAGE 1/2 · 32 TOTAL
  1. TOOL · CL_185430 ·

    Sparse few-shot language model for Bengali achieves 90% sparsity

    Researchers have developed BnBERT-iPET, a novel approach to sparse few-shot language modeling specifically for Bengali. This method utilizes lottery ticket pruning to achieve 90% sparsity, significantly reducing computa…

  2. TOOL · CL_183104 ·

    OpenAI's Privacy Filter shows mixed results in PII detection study

    A new research paper introduces the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5 billion parameter PII detector. The study found that OPF performs well on structured PII types like emails and phon…

  3. TOOL · CL_180469 ·

    Kazakh-Russian code-switching identification bottleneck is annotation, not model

    A new paper on arXiv explores the identification of code-switching between Kazakh and Russian languages, finding that the annotation boundary is more critical than the model used. Researchers developed a gold LID (Langu…

  4. TOOL · CL_167549 ·

    BioSentinel uses XLM-RoBERTa for sexism intent classification in memes

    The BioSentinel team developed a text-centric approach for classifying sexist intent in memes for the EXIST 2026 competition. Their system, built on XLM-RoBERTa-base, utilized a combined loss function incorporating KL d…

  5. RESEARCH · CL_167424 ·

    Indian languages face 8x "tokenizer tax" in LLMs due to English-centric training

    A new research paper highlights a significant disadvantage faced by Indian languages when processed by large language models due to subword tokenization. These tokenizers, primarily trained on English data, result in an…

  6. TOOL · CL_154622 ·

    New dataset and multimodal model tackle Bengali YouTube clickbait

    Researchers have introduced BanClickThumb, a new multimodal dataset designed to detect clickbait in Bengali YouTube videos. The dataset comprises 7,147 thumbnail-title pairs and was used to benchmark various detection m…

  7. TOOL · CL_154405 ·

    New study compares AI strategies for multilingual polarization detection

    Researchers have conducted a comparative study on multilingual polarization detection across 22 languages for SemEval-2026 Task 9. The study evaluated generalist models, language-specific specialists, and ensemble strat…

  8. TOOL · CL_147978 ·

    Urdu fake news detection hindered by dataset length confound

    Researchers have conducted a study on Urdu fake news detection, highlighting significant challenges in cross-dataset generalization. Using the XLM-RoBERTa model and two distinct Urdu datasets, the study found that while…

  9. RESEARCH · CL_147780 ·

    New VEXMLM model boosts performance for African languages

    Researchers have developed VEXMLM, a new language model designed to improve performance on low-resource African languages, specifically Amharic and Tigrinya, which use the Ge'ez script. This model addresses issues of hi…

  10. TOOL · CL_141554 ·

    New benchmark destroR tests and defends Bangla NLP models against attacks

    Researchers have introduced destroR, a new pipeline designed to evaluate and enhance the adversarial robustness of Bangla language transfer models. The system includes three novel meaning-preserving attack methods: a pa…

  11. TOOL · CL_128752 ·

    New NLP framework predicts fake news and mob violence

    Researchers have developed a multimodal Natural Language Processing (NLP) framework designed to detect fake news and predict violence-driven mob activity. This system integrates text and visual data, utilizing XLM-RoBER…

  12. TOOL · CL_117824 ·

    New benchmark reveals AI detectors fail on non-Standard American English dialects

    A new benchmark, DIA-HARM, has been introduced to evaluate the performance of harmful content detection models across 50 English dialects. Researchers found that these models, predominantly trained on Standard American …

  13. RESEARCH · CL_119603 ·

    LLMs show greater robustness in noisy Bangla text event detection than encoder models

    A new research paper evaluates the robustness of different AI model architectures for event detection in noisy Bangla text. The study found that while encoder-only models like BanglaBERT and XLM-R perform better on clea…

  14. RESEARCH · CL_116074 ·

    Small LLMs rival frontier models in relation extraction tasks

    A new research paper explores the effectiveness of large language models (LLMs) for cross-lingual relation extraction, specifically focusing on Romanian. The study found that while LLMs like Gemma 4 31B show a performan…

  15. RESEARCH · CL_84458 ·

    New datasets and model advance emotional validation in AI dialogue

    Researchers have introduced M-EDESConv and M-TESC, new multilingual datasets for emotional validation in dialogue systems, supporting tasks like response identification and timing detection. They also propose MEGUMI, a …

  16. TOOL · CL_65867 ·

    New SindBERT model advances Turkish NLP capabilities

    Researchers have developed SindBERT, a new large-scale RoBERTa-based language model specifically for Turkish. Trained on over 300 GB of Turkish text, SindBERT is available in base and large configurations, marking the f…

  17. RESEARCH · CL_56687 ·

    Perplexity AI open-sources Rust tokenizer, slashing LLM inference latency

    Perplexity AI has open-sourced a new Unigram tokenizer implemented in Rust, which significantly reduces latency and CPU utilization in LLM inference. This new tokenizer achieves up to a 5x lower p50 latency compared to …

  18. TOOL · CL_54988 ·

    Perplexity AI open-sources Unigram tokenizer for 5x speedup

    Perplexity AI has open-sourced a new Unigram tokenizer designed to significantly improve CPU performance. This new tokenizer achieves a 5x reduction in latency compared to HuggingFace's implementation and a 2x reduction…

  19. TOOL · CL_51316 ·

    New dataset boosts Persian social media text classification

    Researchers have introduced PerSoMed, a new large-scale dataset designed for classifying Persian social media text. The dataset contains 36,000 posts across nine categories, with each category having 4,000 samples to en…

  20. TOOL · CL_50954 ·

    AI models struggle with evolving legal language across geopolitical shifts

    Researchers investigated temporal concept drift in legal judgment prediction by training transformer models on Ukrainian court decisions from different geopolitical eras. They found that models trained on older data per…