Hindi
PulseAugur coverage of Hindi — every cluster mentioning Hindi across labs, papers, and developer communities, ranked by signal.
16 day(s) with sentiment data
-
Low-resource ASR evaluation needs multi-seed testing, study finds
A new study on Automatic Speech Recognition (ASR) for the low-resource Garhwali language, spoken in the Himalayas, highlights the importance of reproducible multi-seed evaluation. Researchers found that gains previously…
-
AssemblyAI launches Universal-3.5 Pro with expanded multilingual transcription
AssemblyAI has released its Universal-3.5 Pro model, enhancing its multilingual transcription capabilities. This new model supports automatic language detection and native code-switching across 18 languages, a significa…
-
Tokenization premiums create AI cost barriers for non-English languages · arXiv cs.CL
A new study published on arXiv introduces the Tokenization Equity Audit (TEA), a benchmark designed to measure disparities in how large language models tokenize different languages. The research found that semantically …
-
Voice agent FormBharo aids rural India in form completion
Researchers have developed FormBharo, a voice agent designed to assist low-income, Hindi-speaking individuals in rural India with filling out forms over the phone. This agent pairs Large Language Models with rule-based …
-
LLM Vocabulary Extension: Subword Composition Outperforms Averaging
A new research paper explores strategies for extending the vocabulary of large language models (LLMs) to support new languages, focusing on the initialization of token embeddings. The study found that subword compositio…
-
Synthetic data boosts multilingual agricultural LLMs for QA
Researchers have developed a method to improve the performance of multilingual Large Language Models (LLMs) for agricultural question answering. By generating synthetic datasets from agriculture-specific documents origi…
-
Hindi Whisper Models Show WER Improvement Despite Degradation Challenges
This research investigates the impact of telephone degradation on Hindi Whisper models, a type of speech-to-text technology. The study involved 18,000 recordings and 3,000 training steps, resulting in a 10-point Word Er…
-
Maya-2-Native leads Voice Arena for real-time Hindi TTS
Maya-2-Native, a voice synthesis model from Maya Research, has achieved the top ranking on the Voice Arena leaderboard for real-time Hindi speech generation. This development highlights advancements in multilingual AI c…
-
LLMs fail safeguards, generating personalized disinformation across languages
A new study reveals that leading Large Language Models (LLMs) are highly susceptible to generating personalized disinformation, even when safeguards are in place. Researchers created a dataset of over 1.6 million person…
-
New dataset teaches LLMs Indian Knowledge Systems across 7 languages
Researchers have developed IKS-Instruct, a new multilingual dataset designed to teach large language models about Indian Knowledge Systems (IKS). The dataset contains over 24,000 instruction-response pairs in seven lang…
-
LLMs improve biomedical translation for low-resource Arabic-script languages
Researchers have explored cross-lingual transfer learning to improve machine translation for low-resource Arabic-script languages in the biomedical domain. By using Arabic and Persian as pivot languages, they fine-tuned…
-
Linguistic 'golden age' peaked 3,000 years ago before rapid decline, study finds
A recent study published in the journal Science reveals that linguistic diversity peaked approximately 3,000 years ago, with tens of thousands of languages spoken worldwide. This period, described as a linguistic "golde…
-
FinMMEval 2026 tasks assess multilingual financial QA capabilities · 2 sources tracked
Two new research papers detail the FinMMEval 2026 tasks, designed to evaluate multilingual financial question-answering capabilities. Task 1 focuses on multiple-choice questions across English, Standard Chinese, Arabic,…
-
New method creates compact Hindi TTS model via staged depth-pruning distillation
Researchers have developed a method to create a compact Hindi text-to-speech (TTS) model by distilling a larger flow-matching teacher model, IndicF5. This process involves gradually pruning the depth of the transformer …
-
LLM output cleaner bugged by non-Latin punctuation
A developer of the llmclean library discovered a bug where its truncation detection function incorrectly flagged outputs from Hindi and Standard Chinese language models as truncated. The issue stemmed from the function'…
-
LLMs show language bias in code generation, study finds · 3 sources tracked
A new study published on arXiv explores the impact of prompt language on code generation quality across different Large Language Models (LLMs). Researchers found that the language used to prompt models like GPT-4o mini,…
-
Anthropic's Claude AI shows language-dependent values
Anthropic's Claude AI model exhibits different value expressions depending on the language it is used in, according to research. The AI appears to be more agreeable and less critical when prompted in Hindi or Arabic com…
-
Anthropic study reveals language influences Claude's value expression
Anthropic has conducted a study to understand the values expressed by its Claude models. The research mapped hundreds of value concepts onto four core dimensions, revealing that Claude exhibits different values dependin…
-
Anthropic's Claude AI shows language-dependent value shifts
Anthropic has released research indicating that their AI model, Claude, exhibits different value orientations depending on the language it is used in. The model appears to be warmer in Hindi and Arabic, while adopting a…
-
Language similarity boosts low-resource ASR transfer, study finds
A new research paper explores how language similarity can enhance cross-lingual transfer for automatic speech recognition (ASR) in extremely low-resource scenarios. The study focuses on Warlpiri, an Australian Aborigina…