Slovak
PulseAugur coverage of Slovak — every cluster mentioning Slovak across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
LLMs show weak context-memory conflict on counterfactual data, study finds
A new research paper explores how the perceived plausibility of input data affects the faithfulness of large language models (LLMs). The study generated text in multiple languages, including low-resource ones like Czech…
-
Multilingual NLP models detect harmful and verifiable social media posts
Researchers have developed multilingual transformer-based NLP models capable of detecting social media posts that contain verifiable factual claims and harmful content. This study involved dataset collection, pre-proces…
-
Unicode CLDR defines 1-6 plural forms for language localization
The number of plural forms a language requires for software localization can range from one to six, as defined by the Unicode Consortium's Common Locale Data Repository (CLDR). This count is not a linguistic absolute bu…
-
New Slovak Text Embedding Benchmark and Models Released
Researchers have introduced SkMTEB, a new benchmark designed to evaluate text embedding models specifically for the Slovak language. This benchmark includes 31 datasets across 7 task types, significantly expanding cover…
-
New dataset challenges language ID systems with cousin languages and noise
Researchers have introduced CHALIS, a new dataset designed to test language identification systems in challenging scenarios. The dataset includes examples of closely related languages and text with orthographic noise, s…