A Thai research team has developed a new dataset called Mangosteen, comprising 47 billion tokens, specifically for training Thai large language models (LLMs). They utilized the Dolma data-curation toolkit, originally developed by the Allen Institute for Artificial Intelligence, to filter existing web datasets. This process resulted in a more focused corpus that enhanced the performance of Thai LLMs, even with a reduced amount of data compared to broader datasets. AI
IMPACT This development could lead to more capable and specialized Thai language models, improving AI performance for the region.
RANK_REASON Research team develops a new dataset for LLMs using an existing toolkit. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Bluesky Jetstream — AI desk →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →