Two new research papers explore the mechanisms and training of diffusion large language models (dLLMs). The first paper introduces Trace-Based On-Policy Distillation (TOPD), a framework that improves dLLMs' reasoning abilities by supervising them on their own denoising trajectories, achieving competitive accuracy on mathematical benchmarks with significantly reduced compute. The second paper provides a mechanistic analysis of in-context learning in dLLMs, revealing a bidirectional induction circuit that leverages both past and future context, outperforming autoregressive models when bidirectional context is available. AI
IMPACT These papers advance the understanding and training of diffusion language models, potentially leading to more capable and efficient models for complex tasks like reasoning.
RANK_REASON Two academic papers published on arXiv detailing new methods and analyses for diffusion language models.
- absorbing-mask DLMs
- AR Models
- arXiv
- Diffusion language models
- DLMs
- Hugging Face
- Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models
- transformers
- Diffusion large language models
- dLLMs
- MATH500
- SDAR-4B-Chat
- Trace-Based On-Policy Distillation
- TraDo-4B-Instruct
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →