Bangla
PulseAugur coverage of Bangla — every cluster mentioning Bangla across labs, papers, and developer communities, ranked by signal.
11 day(s) with sentiment data
-
Bangla Math Reasoning Study: CoT Supervision Benefits Vary by Model Strength
A new study, "MathShikkha," investigated the effectiveness of Chain-of-Thought (CoT) supervision for improving mathematical reasoning in small language models (SLMs) specifically for the Bangla language. The research co…
-
Tokenization premiums create AI cost barriers for non-English languages · arXiv cs.CL
A new study published on arXiv introduces the Tokenization Equity Audit (TEA), a benchmark designed to measure disparities in how large language models tokenize different languages. The research found that semantically …
-
New Bangla Sign Language recognition model optimized for mobile deployment
Researchers have developed a new system for recognizing Bangla Sign Language (BdSL) that is designed for deployment on personal devices. The system includes a dataset of over 10,000 expert-validated images of BdSL hand …
-
Sparse few-shot language model for Bengali achieves 90% sparsity
Researchers have developed BnBERT-iPET, a novel approach to sparse few-shot language modeling specifically for Bengali. This method utilizes lottery ticket pruning to achieve 90% sparsity, significantly reducing computa…
-
New BanglaWild benchmark evaluates Bengali scene text recognition for OCR and VLMs
Researchers have introduced BanglaWild, a new benchmark designed to evaluate Bengali scene text recognition for both optical character recognition (OCR) systems and vision-language models (VLMs). The benchmark consists …
-
LLM Safety Alignment Fails Low-Resource Bangla Derogatory Speech
A new research paper published on arXiv investigates the safety alignment of large language models (LLMs) when processing low-resource languages, specifically focusing on derogatory speech in Bangla. The study found tha…
-
New Bengali Sentiment Analysis Framework Employs Continual Learning and LoRA
Researchers have developed SentiBanglaBERT, a novel two-stage framework for sentiment classification in Bengali, a low-resource language. This approach utilizes domain-adaptive continual pretraining and parameter-effici…
-
New Bengali Math Word Problem Dataset Released
Researchers have introduced PatiGonit22K, a new dataset designed to advance the understanding and solving of complex mathematical word problems in Bengali. This dataset expands upon the original PatiGonit dataset, now f…
-
New framework detects political intent in Bengali memes using multimodal fusion
Researchers have developed a novel framework for interpreting multimodal content, specifically focusing on detecting political intent in Bengali memes. This approach utilizes a Vision-Language Model to extract text from…
-
New Bangla-focused benchmark tests MLLMs on document splitting
Researchers have introduced Khondo, a new benchmark designed to evaluate multimodal large language models (MLLMs) on the task of splitting document packets into their constituent parts. This benchmark is unique as it fo…
-
New dataset and multimodal model tackle Bengali YouTube clickbait
Researchers have introduced BanClickThumb, a new multimodal dataset designed to detect clickbait in Bengali YouTube videos. The dataset comprises 7,147 thumbnail-title pairs and was used to benchmark various detection m…
-
New benchmark dataset tackles cultural bias in Bangla language models
Researchers have developed a new benchmark dataset called Culturally Entangled Homograph (CEH) to address the challenge of low-resource language models understanding culturally specific nuances in Bangla. The dataset co…
-
Audio separation harms zero-shot ASR performance, study finds
A new research paper investigates the counterintuitive finding that audio separation can degrade the performance of zero-shot Automatic Speech Recognition (ASR) systems. The study evaluated SAM-Audio as a preprocessing …
-
New framework uses SLMs to combat health misinformation in Bangla
Researchers have developed a novel framework to detect health misinformation in low-resource languages, using Bangla as a case study. The framework integrates Small Language Models (SLMs) with a culturally sensitive Res…
-
New benchmark destroR tests and defends Bangla NLP models against attacks
Researchers have introduced destroR, a new pipeline designed to evaluate and enhance the adversarial robustness of Bangla language transfer models. The system includes three novel meaning-preserving attack methods: a pa…
-
Less common languages can bypass AI safety features, researchers find
Large language models (LLMs) can be more easily tricked or bypassed by using less common languages for prompts, as opposed to English. This is because most LLMs are primarily trained on vast amounts of English-language …
-
New method adapts lightweight ASR models for Bengali language
Researchers have developed a novel method to adapt lightweight speech recognition models, like Moonshine, for morphologically rich languages such as Bengali. The core issue identified was an English-centric tokenizer th…
-
New dataset and model tackle explanatory evidence detection in Bengali memes
Researchers have introduced MemeEvidenceDetect, a novel task focused on identifying explanatory sentences within memes, particularly for the Bangla language. To support this, they developed BanglaMemeEvidence, a dataset…
-
New NLP framework predicts fake news and mob violence
Researchers have developed a multimodal Natural Language Processing (NLP) framework designed to detect fake news and predict violence-driven mob activity. This system integrates text and visual data, utilizing XLM-RoBER…
-
New benchmark targets MLLM comprehension of complex Bangla documents
Researchers have introduced BaFCo, a new benchmark dataset designed to improve document comprehension for Multimodal Large Language Models (MLLMs) in the Bangla language. The dataset consists of 200 complex Bangladeshi …