PulseAugur
EN
LIVE 18:40:19
ENTITY multilingual-BERT

multilingual-BERT

PulseAugur coverage of multilingual-BERT — every cluster mentioning multilingual-BERT across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
19
19 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
19
19 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 20 TOTAL
  1. TOOL · CL_286817 ·

    QuanLing framework quantifies language distance in Western Romance languages

    Researchers have extended the QuanLing framework, which uses pretrained language models to quantify language distance, to Western Romance languages. This framework combines sentence embedding distances, tokenization fra…

  2. RESEARCH · CL_257005 ·

    AI legal assistants developed for Nepal to improve access to justice · 2 sources tracked

    Two research papers introduce AI-powered legal assistants for Nepal, aiming to improve access to justice. The first, NepKANUN, utilizes a retrieval-augmented generation (RAG) framework with a fine-tuned large language m…

  3. RESEARCH · CL_229094 ·

    Hugging Face unveils efficient multimodal encoder NeoMME, study favors encoders for Indic NER

    Hugging Face has introduced NeoMME, a new family of multilingual multimodal encoders designed for efficiency. Unlike many generative models, NeoMME uses a single bidirectional Transformer to process both text and image …

  4. TOOL · CL_223219 ·

    New TabuLM model enhances low-resource language processing with tabular data

    Researchers have developed TabuLM, a novel language model specifically pre-trained on Kinyarwanda tabular data to address the scarcity of resources for low-resource languages. This model extends KinyaBERT-large by incor…

  5. TOOL · CL_223114 ·

    Prompt compression struggles with non-English languages, study finds

    A new study published on arXiv titled "Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors" investigates the effectiveness of prompt compression techniques across different languages. …

  6. TOOL · CL_227836 ·

    TabuLM: First Language Model Pre-trained on Kinyarwanda Tabular Data

    Researchers have developed TabuLM, a novel language model specifically pre-trained on tabular data for Kinyarwanda, a low-resource Bantu language spoken in Rwanda. This model enhances KinyaBERT-large with new embeddings…

  7. TOOL · CL_218155 ·

    New Corpus and Transformer Model Detect Metaphors in Hindi Legal Texts

    Researchers have developed a new method for detecting metaphors in Hindi legal documents, a task previously unaddressed due to a lack of annotated data for low-resource languages. They created the Hindi Legal Metaphor C…

  8. TOOL · CL_212065 ·

    New Nepali-English benchmark for misinformation detection released

    Researchers have developed NepOOC-M, the first publicly available benchmark for detecting out-of-context (OOC) misinformation in Nepali and English. The dataset includes 1,090 image-caption pairs annotated across five t…

  9. TOOL · CL_210414 ·

    New multilingual model NE-BERT boosts NLP for 9 Northeast Indian languages

    Researchers have developed NE-BERT, a new multilingual language model specifically designed for nine underrepresented Northeast Indian languages. This model, trained on approximately 8.3 million sentences, significantly…

  10. TOOL · CL_183255 ·

    New HomoEnsNER model boosts Gujarati NER performance

    Researchers have developed HomoEnsNER, a novel approach to Named Entity Recognition (NER) for the Gujarati language. This method utilizes a homogeneous ensemble of five independently fine-tuned GujaratiBERT models, whic…

  11. TOOL · CL_227984 ·

    Prompt compression fails non-English languages, new paper finds

    A new paper titled "Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors" reveals that prompt compression techniques, designed to reduce LLM inference costs by removing low-information …

  12. TOOL · CL_111167 ·

    Brain's bilingual processing mirrors LLM vector space isomorphism

    Scientists have discovered that the human brain organizes words in bilingual individuals using a geometric map, similar to how Large Language Models (LLMs) utilize vector space isomorphism. This study, published in Cell…

  13. RESEARCH · CL_115257 ·

    HSA_CORAL's GPT-4.1 Mini leads FinCausal 2026 financial causality task

    A research paper details HSA_CORAL's approach to the FinCausal 2026 shared task, focusing on extracting cause-effect relationships from financial texts. The team explored three model families: multilingual BERT for toke…

  14. TOOL · CL_114347 ·

    New method ROMEVA improves Roman Urdu language model adaptation

    A new research paper introduces ROMEVA, a method for expanding the vocabulary of multilingual language models like mBERT to better handle morphologically inconsistent languages such as Roman Urdu. Roman Urdu's inconsist…

  15. TOOL · CL_104744 ·

    New method ROMEVA improves Roman Urdu language model vocabulary

    Researchers have developed ROMEVA, a novel method for expanding the vocabulary of multilingual language models like mBERT to better handle languages with inconsistent spelling, such as Roman Urdu. This approach combines…

  16. TOOL · CL_93525 ·

    Open Diachronic Greek Treebank Released with Indo-European Parallels

    Researchers have developed AthDGC, a comprehensive, open-source dataset and workflow for dependency parsing of the Greek language across eight historical periods. This project, built upon the PROIEL Treebank Family sche…

  17. RESEARCH · CL_48842 ·

    New pipeline creates NLP resource for historical Greek parliamentary text

    Researchers have developed a new, reproducible pipeline for creating a Universal Dependencies-style parsing resource for Katharevousa Greek parliamentary text. This workflow addresses the limitations of current NLP tool…

  18. TOOL · CL_22204 ·

    Bengali AI models show identity biases despite similar data, study finds

    A new paper investigates biases in sentiment analysis models for the Bengali language, a low-resource context. Researchers audited models like mBERT and BanglaBERT, fine-tuned on Bengali sentiment analysis datasets, and…

  19. RESEARCH · CL_20602 ·

    New benchmark study explores neural network performance on Tajik POS tagging

    This paper introduces the first benchmark for part-of-speech tagging in the Tajik language, evaluating various neural network architectures. The study utilized the TajPersParallel corpus, focusing on context-independent…

  20. TOOL · CL_15858 ·

    New Sindhi figurative language dataset SiNFluD released with XLM-RoBERTa-XL benchmark

    Researchers have developed SiNFluD, a new dataset for classifying figurative language in Sindhi. The dataset was compiled from various online sources and annotated by native speakers, achieving a high inter-annotator ag…