Sinhala
PulseAugur coverage of Sinhala — every cluster mentioning Sinhala across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
Sinhala language semantic change analyzed with Llama-3.1-8B
Researchers have developed a computational framework to analyze the diachronic semantic change in the Sinhala language, spanning from the 13th to the 20th century. The study utilized both static embeddings (Word2Vec, Fa…
-
SinLlama model enhances Llama 3-8B for Sinhala language tasks
Researchers have developed SinLlama, a new open-source large language model specifically designed for the Sinhala language. By enhancing the Llama-3-8B model with Sinhala vocabulary and training it on a substantial Sinh…
-
New HelaBERT models boost Sinhala language understanding
Researchers have developed HelaBERT, a new family of BERT-based language models specifically designed to enhance understanding of the Sinhala language. These models, HelaBERT-Small and HelaBERT-Large, were pre-trained o…
-
New framework enables trilingual topic modeling of Sri Lankan parliamentary debates
Researchers have developed a novel framework to perform topic modeling on trilingual parliamentary debates from Sri Lanka, encompassing Sinhala, Tamil, and English. This system overcomes challenges posed by complex PDF …
-
Sinhala-Tamil CLIR research favors embedding models over translation
A new research paper evaluates cross-lingual information retrieval (CLIR) methods for accessing English government information using Sinhala and Tamil queries. The study compared query translation techniques, including …
-
New Python library grapheme-kit improves multilingual NLP metrics
A new open-source Python library called grapheme-kit has been developed to address limitations in existing text processing metrics. These metrics, which typically operate on Unicode code points, can inaccurately represe…
-
New LKValues resource aligns LLMs with Sri Lankan societal values
Researchers have developed LKValues, a new resource suite designed to align large language models (LLMs) with the specific societal values of Sri Lanka. Existing benchmarks often reflect Western norms, failing to captur…
-
New Sinhala dataset launched for Aspect-Based Sentiment Analysis
Researchers have introduced SalAngaBhava, a new dataset designed for Aspect-Based Sentiment Analysis (ABSA) in Sinhala, a low-resource language spoken primarily in Sri Lanka. This dataset comprises Sinhala product revie…
-
Cross-lingual transfer learning shows mixed results for low-resource ASR
Researchers have explored cross-lingual transfer learning to improve automatic speech recognition (ASR) for low-resource languages. One study successfully used Sinhala to enhance Dhivehi ASR, achieving a 12.89% Word Err…
-
New Sinhala OCR Dataset and Model Achieve State-of-the-Art Performance
Researchers have developed a new dataset, sinhala-ocr-lk-acts-1010, to improve Optical Character Recognition (OCR) for the Sinhala language, which is spoken by approximately 16 million people in Sri Lanka. This dataset …
-
New Sinhala OCR Dataset and LightOnOCR-2-1B Achieve State-of-the-Art Performance
Researchers have developed a new dataset, sinhala-ocr-lk-acts-1010, to improve Optical Character Recognition (OCR) for Sinhala, a language spoken by approximately 16 million people. This dataset comprises 1,010 page-lev…
-
Sinhala NLP research hub releases transliteration systems and data resources
A new paper introduces the Swa-bhasha Resource Hub, a collection of data and algorithms for Romanized Sinhala to Sinhala transliteration developed between 2020 and 2025. These resources are crucial for advancing Sinhala…