Researchers have introduced TabiBERT, a new large-scale foundation model for the Turkish language based on the ModernBERT architecture. This model was trained on a massive dataset of over one trillion tokens, encompassing web text, scientific publications, and source code, and supports a context length of 8,192 tokens. To evaluate TabiBERT, a new benchmark called TabiBench was developed, comprising 27 datasets across eight task categories. TabiBERT demonstrated superior performance compared to previous Turkish models, particularly in question answering and code retrieval tasks. AI
IMPACT Establishes a new state-of-the-art for Turkish language models and provides a standardized evaluation framework for future research.
RANK_REASON The cluster describes a new research paper introducing a novel language model and benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →