Kazakh
PulseAugur coverage of Kazakh — every cluster mentioning Kazakh across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New research tackles AI hallucinations in video and language models
Researchers are developing new methods to combat hallucinations in AI models, particularly in video-language and large language models. One approach, CounterVid, uses counterfactual video generation to create synthetic …
-
Kazakh-Russian code-switching identification bottleneck is annotation, not model
A new paper on arXiv explores the identification of code-switching between Kazakh and Russian languages, finding that the annotation boundary is more critical than the model used. Researchers developed a gold LID (Langu…
-
Cross-lingual transfer in Turkic languages shows strong pair-specific performance
Researchers have investigated cross-lingual transfer techniques for machine translation within the Turkic language family, focusing on Turkish, Azerbaijani, Uzbek, Kazakh, and Kyrgyz. Their findings indicate that transf…
-
ai-sage releases GigaAM Multilingual speech models
ai-sage has released GigaAM Multilingual, a family of Conformer-based foundation models. These models, available in 220M and 600M parameter variants, have been pre-trained on over 2 million hours of speech data spanning…
-
Study finds green color names most stable across languages
A new study published on arXiv investigates the stability of color naming across different languages and shades. Researchers used a free color-naming experiment with 92 participants who assigned names to red, yellow, an…
-
New Kazakh Prompt Dataset Reveals LLM Safety Gaps
Researchers have developed KZ-SafetyPrompts, a new dataset designed to evaluate the safety of large language models (LLMs) in the Kazakh language. The dataset comprises 5,717 prompts across eleven risk categories, inclu…