Researchers have developed GrowMTP, a novel method that trains a draft head for speculative decoding entirely within the reinforcement learning (RL) loop. This approach eliminates the need for pre-training draft heads separately, reducing training costs. GrowMTP leverages the narrow rollout distribution and continuous supervision signals generated during RL training to optimize the draft head. Experiments on Qwen3-4B, MiMo-7B-SFT, and Qwen3.5-4B-Base models demonstrated significant speedups in both rollout generation and end-to-end training times. AI
IMPACT Accelerates LLM training by integrating draft head optimization directly into the RL process, reducing computational overhead.
RANK_REASON The cluster describes a new research paper detailing a novel method for training AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- GrowMTP
- Hugging Face
- large-language models
- MiMo-7B-SFT
- Qwen3-4B
- Qwen3.5-4B-Base
- reinforcement learning
- speculative decoding
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →