Two new research papers introduce novel methods for accelerating the inference speed of large language models. The first paper, "ReTrace," proposes a technique that conditions each draft block on the rejected suffix from the previous round, improving average acceptance length and decoding speed for models like Qwen3. The second paper, "Trajectory-Level Speculative Decoding," presents a framework for diffusion-based language models (dLLMs) that speculates over denoising trajectories, achieving significant speedups over vanilla dLLMs and existing frameworks like Fast-dLLM++. AI
IMPACT These advancements in speculative decoding could lead to faster and more efficient deployment of large language models across various applications.
RANK_REASON Two academic papers published on arXiv introducing new methods for LLM inference acceleration.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →