PulseAugur
EN
LIVE 06:28:50

New Turkish LLM MoganBert-TR trained with CLM-to-MLM curriculum

Researchers have developed MoganBert-TR, a new Turkish encoder foundation model, and its accompanying embedding model, MoganBert-Embed. Trained from scratch on a filtered Turkish corpus using a novel CLM-to-MLM curriculum, MoganBert-TR demonstrates significant improvements over traditional MLM approaches, particularly in retrieval tasks. The model achieves state-of-the-art results on benchmarks like TrGLUE and TabiBench, while MoganBert-Embed excels in embedding performance, ranking first on MTEB(Turkish) despite its smaller size. AI

IMPACT Introduces a new training curriculum that improves performance on Turkish language tasks and offers a more efficient embedding model.

RANK_REASON The cluster describes a new academic paper detailing the creation and evaluation of a novel language model. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New Turkish LLM MoganBert-TR trained with CLM-to-MLM curriculum

How we ranked this

Signal score
30 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster describes a new academic paper detailing the creation and evaluation of a novel language model. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay ·

    MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum

    arXiv:2608.25768v1 Announce Type: new Abstract: Turkish encoder models have adopted modern architectures while leaving the pretraining objective fixed at masked language modelling. This paper introduces MoganBert-TR, a 149M-parameter Turkish encoder foundation model trained from …