Researchers have developed MoganBert-TR, a new Turkish encoder foundation model, and its accompanying embedding model, MoganBert-Embed. Trained from scratch on a filtered Turkish corpus using a novel CLM-to-MLM curriculum, MoganBert-TR demonstrates significant improvements over traditional MLM approaches, particularly in retrieval tasks. The model achieves state-of-the-art results on benchmarks like TrGLUE and TabiBench, while MoganBert-Embed excels in embedding performance, ranking first on MTEB(Turkish) despite its smaller size. AI
IMPACT Introduces a new training curriculum that improves performance on Turkish language tasks and offers a more efficient embedding model.
RANK_REASON The cluster describes a new academic paper detailing the creation and evaluation of a novel language model. [lever_c_demoted from research: ic=1 ai=1.0]
- CLM-to-MLM
- Hugging Face
- MoganBert-Embed
- MoganBert-TR
- MS MARCO
- MTEB(Turkish)
- TabiBench
- TabiBERT
- TrGLUE
- Turkish
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →