Researchers have discovered that masked diffusion large language models (LLMs) can leverage end-of-sequence (EoS) tokens for hidden reasoning, enhancing their performance on complex tasks. By padding answers with EoS tokens beyond the correct length, these models appear to utilize the representations of these tokens as additional computational capacity. Experiments with models like LLaDA1.5 and Dream-v0 on tasks such as Addition, Entity Tracking, and Sudoku confirmed that adding EoS tokens improves accuracy. Further causal interventions, where hidden states of EoS tokens were transferred, increased the likelihood of counterfactual answers, indicating latent reasoning capabilities within these tokens. AI
IMPACT This research suggests a new method for enhancing LLM reasoning capabilities by leveraging EoS tokens, potentially improving performance on complex tasks without architectural changes.
RANK_REASON The cluster contains a research paper detailing a novel finding about LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- Addition
- Dream-v0
- EoS tokens
- GSM8K
- LLaDA1.5
- LLaDA2.0-mini
- Masked diffusion LLMs
- Sarah Breckner
- sudoku
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →