Researchers have identified a significant bottleneck in extending language models (LMs) with new vocabulary for domain-specific tasks, such as generative recommendation. The standard method of initializing new tokens with the mean of existing embeddings leads to a collapse of distinctions that fine-tuning struggles to recover. A new approach, Grounded Token Initialization (GTI), proposes mapping novel tokens to distinct, semantically meaningful locations in the pretrained embedding space before fine-tuning. This lightweight grounding stage has demonstrated superior performance across multiple generative recommendation benchmarks compared to mean initialization and other adaptation methods, suggesting that initialization quality is crucial for effective vocabulary extension. AI
IMPACT This research could lead to more effective domain-specific language models, particularly in recommendation systems, by improving how new concepts are integrated.
RANK_REASON Academic paper detailing a new method for language model vocabulary extension. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Daiwei Chen
- Generative recommendation
- Grounded Token Initialization
- Hugging Face
- Language models
- Semantic ID
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →