Malayalam
PulseAugur coverage of Malayalam — every cluster mentioning Malayalam across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
Srijika system generates OpenType fonts for nine Indic scripts
Researchers have developed Srijika, a novel system designed to create installable OpenType fonts for nine Indic scripts, including Devanagari, Tamil, and Bengali. Instead of generating fonts from scratch, Srijika restyl…
-
New benchmarks reveal LLM struggles with Indic languages and translation mechanics · 3 sources tracked
Researchers have developed VakyArth, a new benchmark designed to evaluate the pragmatic competence of large language models (LLMs) specifically within Indic languages like Hindi, Punjabi, Tamil, and Malayalam. Initial f…
-
New benchmark released for Indic language quality estimation and post-editing
Researchers have introduced IndicQE-APE, a new benchmark designed to consolidate and evaluate quality estimation and automatic post-editing for Indic languages. This benchmark combines data from WMT shared tasks and an …
-
Tamil and Malayalam OCR struggles with complex vowel sign placement
Optical character recognition (OCR) for Tamil and Malayalam scripts faces challenges primarily with vowel signs due to their placement and complex interactions with consonants. These scripts are abugidas, where vowel si…
-
Monolingual models outperform multilingual on Dravidian languages
Researchers have developed and evaluated five GPT-2 architecture models to assess the performance of multilingual language models on Dravidian languages. Four of these models were trained monolingually for Tamil, Telugu…
-
New dataset teaches LLMs Indian Knowledge Systems across 7 languages
Researchers have developed IKS-Instruct, a new multilingual dataset designed to teach large language models about Indian Knowledge Systems (IKS). The dataset contains over 24,000 instruction-response pairs in seven lang…
-
New AI system GeoMVC tackles misogyny in multimodal memes
Researchers have developed a new system called GeoMVC for detecting misogyny in internet memes, a task complicated by the interplay between visual and textual elements and cultural context. The system employs a Geometri…
-
Indian languages face 8x "tokenizer tax" in LLMs due to English-centric training
A new research paper highlights a significant disadvantage faced by Indian languages when processed by large language models due to subword tokenization. These tokenizers, primarily trained on English data, result in an…
-
New benchmark and fine-tuning technique improve Indic language ASR
Researchers have developed Vividh-ASR, a new benchmark designed to evaluate automatic speech recognition (ASR) models on Indic languages, specifically Hindi and Malayalam. This benchmark categorizes audio into four tier…
-
SCRIBE framework improves ASR for Indic languages with new error analysis
Researchers have introduced SCRIBE, a new diagnostic framework designed to improve automatic speech recognition (ASR) for Indic languages. Unlike traditional metrics like Word Error Rate (WER), SCRIBE categorizes errors…
-
New benchmark tackles ASR bias in Indic languages
Researchers have developed Vividh-ASR, a new benchmark designed to evaluate automatic speech recognition (ASR) models for Indic languages, specifically Hindi and Malayalam. This benchmark categorizes audio into four tie…