Tibetan
PulseAugur coverage of Tibetan — every cluster mentioning Tibetan across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
China's DeepZang AI targets Tibetan language tech and ethnic unity
China has developed DeepZang AI, a large language model designed to advance Tibetan language technology and promote ethnic unity. A recent seminar in Hohhot, Inner Mongolia, focused on improving the model's capabilities…
-
UniLipi OCR model unifies 13 Indic scripts for historical manuscripts
Researchers have developed UniLipi, a novel unified multi-script Optical Character Recognition (OCR) model designed for historical Indic manuscripts. This single framework can process 13 different Indic scripts, address…
-
China launches hi-tech initiative to map and preserve linguistic diversity
China has launched a new language resources hall at the Hunan Museum in Changsha, showcasing over 100 ethnic minority languages and Chinese dialects. This initiative, developed over a decade with 1,800 survey sites, aim…
-
AI models fail silently in non-English languages, research shows
AI models often perform poorly in languages other than English, despite passing English-language tests. Research indicates significant accuracy drops in languages like Swahili, Tibetan, and Arabic, with models like GPT-…
-
Research reveals flawed tokenizer design limits multilingual AI models
A new research paper highlights a significant limitation in multilingual tokenizers used by many AI models, including those from Hugging Face and potentially impacting models like GPT-4o. The study identifies that token…
-
New RL Framework Enhances Low-Resource Language Generation
Researchers have developed a new framework called Source-Grounded Semantic Reinforcement Learning (SG-SRL) to improve low-resource target-language generation. This method leverages abundant source-language monolingual d…
-
Xingchen AGI Lab develops first large-model Tibetan TTS system
Researchers have developed Tibetan-TTS, a novel text-to-speech system designed for the Tibetan language, which is characterized by limited data and dialectal variations. This system leverages a large speech synthesis mo…
-
New TTS framework synthesizes Tibetan dialects from limited data
Researchers have developed FMSD-TTS, a novel few-shot text-to-speech system designed to generate speech for the low-resource Tibetan language across its three main dialects: Ü-Tsang, Amdo, and Kham. The system utilizes …