An experiment exploring the impact of vocabulary size on small language models found that capping the vocabulary of a technical corpus (cs.CL abstracts) significantly harmed model performance. While a smaller vocabulary might seem to simplify next-token prediction, this study demonstrated that it leads to a higher rate of unknown tokens and a collapse in novel n-grams, resulting in repetitive and grammatically skeletal output. The findings suggest that the success of the TinyStories dataset, which used a small vocabulary, was due to the inherent simplicity of its domain, not just the vocabulary size itself. AI
IMPACT Demonstrates that vocabulary size is not a simple lever for improving small model performance on complex domains.
RANK_REASON Research paper detailing experimental findings on LLM vocabulary impact. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →