Sanskrit
PulseAugur coverage of Sanskrit — every cluster mentioning Sanskrit across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
LLM translation errors in classical texts assessed without human review
Researchers have developed a novel method for evaluating the accuracy of Large Language Model (LLM) translations of classical texts without requiring human references. The study focused on Pali-to-English translation, c…
-
Sanskrit tokenization penalty higher than English, study finds
A new research paper investigates the tokenization efficiency of Sanskrit compared to English and Hindi when processed by modern language models. The study found that Sanskrit requires significantly more tokens per unit…
-
New NER Benchmark for Classical Sanskrit Developed
Researchers have developed Padārtha, a novel Named Entity Recognition (NER) benchmark specifically designed for classical Sanskrit texts. This benchmark grounds its annotation schema in the Nyāya-Vaiśeṣika ontological s…
-
Gurus embrace AI chatbots for spiritual guidance, meeting demand for AI-driven answers
Gurus are increasingly turning to AI chatbots to offer spiritual guidance, a trend driven by people seeking answers from artificial intelligence. This shift is particularly notable in traditions like Hinduism, where hum…
-
Sanskrit TTS system Vāgdhenu converts meter-aware shlokas to chants
Researchers have developed Vāgdhenu, a novel text-to-speech system designed to convert Sanskrit shlokas into high-fidelity chanted recitations. The system utilizes an existing TTS backbone and neural vocoder, enhanced w…
-
New NLP task targets Sanskrit glossary generation with benchmark
Researchers have introduced "grounded glossary generation," a new NLP task focused on extracting Sanskrit phrases and their meanings from sloka-translation pairs. They developed a benchmark dataset of over 31,000 triple…
-
New OCR method improves Sanskrit manuscript digitization
Researchers have developed an iterative fine-tuning approach for optical character recognition (OCR) to improve the digitization of complex historical Sanskrit manuscripts. This method adapts to manuscript-specific layo…
-
OCR challenges for Thai, Khmer, Korean, and Ethiopic scripts detailed
Optical character recognition (OCR) for scripts like Thai, Khmer, Korean, and Ethiopic presents unique challenges beyond standard Latin-based text. Thai OCR struggles with word segmentation due to the absence of spaces …
-
New AI approach enhances Sanskrit poetry generation with prosody focus
Researchers have developed Pingala, a novel decoding approach for generating Sanskrit poetry that emphasizes prosody and semantic coherence. By segmenting verses into grouped lines and favoring longer tokens, Pingala im…
-
Research questions statistical methods for deciphering Indus script
A new research paper published on arXiv questions the validity of statistical measures used to decipher ancient scripts, particularly the Indus script. The study introduces SIGIL, a generative emblem system designed to …
-
New system SIGIL questions statistical methods for script decipherment
A new study challenges the statistical methods used to decipher unknown scripts, particularly the Indus script. Researchers developed a generative emblem system called SIGIL, which mimics the statistical properties of t…
-
New dataset teaches LLMs Indian Knowledge Systems across 7 languages
Researchers have developed IKS-Instruct, a new multilingual dataset designed to teach large language models about Indian Knowledge Systems (IKS). The dataset contains over 24,000 instruction-response pairs in seven lang…
-
New BHARATI tokenizers boost efficiency for classical Indian languages
Researchers have developed BHARATI, a new set of tokenizers specifically designed for classical Indian languages like Sanskrit and Tamil. Unlike standard algorithms that struggle with the agglutinative morphology and sa…
-
New TTS system Vāgdhenu generates Sanskrit chanting
A new text-to-speech system named Vāgdhenu has been developed to generate Sanskrit chanting. This system aims to accurately reproduce the intonation and pronunciation required for traditional Sanskrit recitations. The p…
-
Lifelong learning dismissed as hoax; single-celled organisms cited as true learners
The concept of lifelong learning is critiqued as a widespread illusion, with traditional education and corporate training often amounting to mere memorization or job-performance rather than genuine skill acquisition. Th…
-
Computational linguistics studies analyze lexical transmission in religious texts · 4 sources tracked
Two new computational linguistics studies analyze lexical transmission and stylistic features within religious texts. The first paper examines Bengali and Sanskrit devotional literature from the 8th to 19th centuries, u…
-
Pāninian grammar offers unified framework for Indic language NLP
A new research paper proposes a unified computational framework for Indic languages, drawing inspiration from Pāninian grammar. The authors argue that the shared morphosyntactic architecture across these languages, form…
-
NoRIN advances time-series forecasting with non-linear normalization
Researchers have introduced NoRIN, a novel non-linear reversible normalization technique for time-series forecasting that goes beyond the linear affine transformations of existing methods like RevIN. NoRIN utilizes a Jo…
-
Researchers create Naamah, a large synthetic Sanskrit NER dataset using LLMs
Researchers have developed Naamah, a synthetic dataset of over 100,000 Sanskrit sentences designed to improve Named Entity Recognition (NER) for classical Sanskrit literature. The dataset was generated by combining enti…