Sinhala
PulseAugur coverage of Sinhala — every cluster mentioning Sinhala across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
New Python library grapheme-kit improves multilingual NLP metrics
A new open-source Python library called grapheme-kit has been developed to address limitations in existing text processing metrics. These metrics, which typically operate on Unicode code points, can inaccurately represe…
-
New LKValues resource aligns LLMs with Sri Lankan societal values
Researchers have developed LKValues, a new resource suite designed to align large language models (LLMs) with the specific societal values of Sri Lanka. Existing benchmarks often reflect Western norms, failing to captur…
-
New Sinhala dataset launched for Aspect-Based Sentiment Analysis
Researchers have introduced SalAngaBhava, a new dataset designed for Aspect-Based Sentiment Analysis (ABSA) in Sinhala, a low-resource language spoken primarily in Sri Lanka. This dataset comprises Sinhala product revie…
-
Cross-lingual transfer learning shows mixed results for low-resource ASR
Researchers have explored cross-lingual transfer learning to improve automatic speech recognition (ASR) for low-resource languages. One study successfully used Sinhala to enhance Dhivehi ASR, achieving a 12.89% Word Err…
-
New Sinhala OCR Dataset and Model Achieve State-of-the-Art Performance
Researchers have developed a new dataset, sinhala-ocr-lk-acts-1010, to improve Optical Character Recognition (OCR) for the Sinhala language, which is spoken by approximately 16 million people in Sri Lanka. This dataset …
-
New Sinhala OCR Dataset and LightOnOCR-2-1B Achieve State-of-the-Art Performance
Researchers have developed a new dataset, sinhala-ocr-lk-acts-1010, to improve Optical Character Recognition (OCR) for Sinhala, a language spoken by approximately 16 million people. This dataset comprises 1,010 page-lev…
-
Sinhala NLP research hub releases transliteration systems and data resources
A new paper introduces the Swa-bhasha Resource Hub, a collection of data and algorithms for Romanized Sinhala to Sinhala transliteration developed between 2020 and 2025. These resources are crucial for advancing Sinhala…