A new research paper explores strategies for extending the vocabulary of large language models (LLMs) to support new languages, focusing on the initialization of token embeddings. The study found that subword composition methods, particularly asymmetric variants, outperform traditional vocabulary averaging and external initialization techniques. The optimal configuration involved initializing the input embedding matrix with uniform subword averaging and Hindi-specific norm calibration, and the output language modeling head with character-length-weighted subword averaging. This approach significantly reduced continued pre-training steps and improved accuracy compared to the standard baseline, while also demonstrating that initialization loss and bits-per-byte are poor predictors of downstream convergence. AI
IMPACT Optimizes LLM adaptation to new languages, potentially reducing training costs and improving multilingual capabilities.
RANK_REASON Academic paper detailing a systematic study of LLM techniques. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →