A new paper proposes a novel perspective on tokenization in language models, arguing that it should be viewed as output supervision rather than just input preprocessing. The researchers conducted experiments demonstrating that output tokenization significantly impacts a model's learning dynamics and internal representations, independent of input tokenization. Their analysis of recent CL papers on numeric reasoning revealed that this crucial aspect of tokenization is often overlooked, with many studies comparing models across different tokenization strategies without acknowledging the resulting differences in supervision. AI
IMPACT This research could lead to more principled comparisons of language models and potentially influence future model design by highlighting the impact of output tokenization.
RANK_REASON The cluster contains a research paper detailing a new theoretical framing and experimental findings in NLP. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- arXivLabs
- CatalyzeX Code Finder for Papers
- CL papers
- Computation and Language
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →