Researchers have developed two new subword segmental language models, SubSegGPT and SubSegDeBERTa, for the 2026 BabyLM Challenge. These models learn tokenization during the pretraining phase, allowing them to discover optimal subword units. SubSegDeBERTa showed notable gains in zero-shot evaluation for the Strict track, while SubSegGPT outperformed tokenization-based baselines on the Strict-small track. The findings suggest that learnable subword tokenization enhances sample-efficiency in language model pretraining. AI
IMPACT Demonstrates potential for improved sample-efficiency in language model pretraining through learned tokenization.
RANK_REASON Academic paper detailing new models and experimental results. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →