Tamil
PulseAugur coverage of Tamil — every cluster mentioning Tamil across labs, papers, and developer communities, ranked by signal.
7 day(s) with sentiment data
-
Monolingual models outperform multilingual on Dravidian languages
Researchers have developed and evaluated five GPT-2 architecture models to assess the performance of multilingual language models on Dravidian languages. Four of these models were trained monolingually for Tamil, Telugu…
-
Tokenization premiums create AI cost barriers for non-English languages · arXiv cs.CL
A new study published on arXiv introduces the Tokenization Equity Audit (TEA), a benchmark designed to measure disparities in how large language models tokenize different languages. The research found that semantically …
-
New MERaLiON-GR model achieves state-of-the-art gender recognition across multiple languages
Researchers have developed MERaLiON-GR, a novel speech gender recognition model capable of classifying gender for both English and several Southeast Asian languages. This model is built upon MERaLiON-SpeechEncoder-2, a …
-
Research questions statistical methods for deciphering Indus script
A new research paper published on arXiv questions the validity of statistical measures used to decipher ancient scripts, particularly the Indus script. The study introduces SIGIL, a generative emblem system designed to …
-
Tamil language models get morphology-aware enhancement for better translation
Researchers have developed a novel morphology-aware system for Tamil language models, enhancing translation capabilities. This system integrates the ThamizhiMorph analyzer and generator with a byte-exact semantic tokeni…
-
New system SIGIL questions statistical methods for script decipherment
A new study challenges the statistical methods used to decipher unknown scripts, particularly the Indus script. Researchers developed a generative emblem system called SIGIL, which mimics the statistical properties of t…
-
New dataset teaches LLMs Indian Knowledge Systems across 7 languages
Researchers have developed IKS-Instruct, a new multilingual dataset designed to teach large language models about Indian Knowledge Systems (IKS). The dataset contains over 24,000 instruction-response pairs in seven lang…
-
New BHARATI tokenizers boost efficiency for classical Indian languages
Researchers have developed BHARATI, a new set of tokenizers specifically designed for classical Indian languages like Sanskrit and Tamil. Unlike standard algorithms that struggle with the agglutinative morphology and sa…
-
New AI system GeoMVC tackles misogyny in multimodal memes
Researchers have developed a new system called GeoMVC for detecting misogyny in internet memes, a task complicated by the interplay between visual and textual elements and cultural context. The system employs a Geometri…
-
Transformer Architecture Built From Scratch for English-Tamil Translation
An individual has developed and trained a Transformer neural network architecture from scratch using PyTorch. This model is specifically designed for English-to-Tamil machine translation and is based on the "Attention I…
-
New Python library grapheme-kit improves multilingual NLP metrics
A new open-source Python library called grapheme-kit has been developed to address limitations in existing text processing metrics. These metrics, which typically operate on Unicode code points, can inaccurately represe…
-
New N-VSSM Model Outperforms Claude Opus 4.5 in Long-Form Narrative Consistency
Researchers have developed NarrativeWorldBench, a new benchmark designed to evaluate large language models (LLMs) on their ability to maintain narrative consistency in long-form audio dramas. Current frontier LLMs strug…
-
New ASR Error Analysis Tool Breaks Script Barriers
Researchers have developed a new automated alignment mechanism designed to improve the analysis of Automatic Speech Recognition (ASR) errors, particularly for languages that do not use the Latin script. This method is l…
-
Linear memory capacity depends on retrieval: $n\log n$ for top-1, $n$ for listwise
Researchers have analyzed the capacity limits of linear associative memory, finding that the retrieval criterion significantly impacts how many associations can be stored. For top-1 retrieval, where a signal must outper…
-
New benchmark evaluates Indic TTS accent fidelity across six dimensions
Researchers have introduced PSP, a new benchmark designed to evaluate the accent accuracy of text-to-speech (TTS) systems for Indic languages. Unlike existing metrics that focus on intelligibility and naturalness, PSP s…