Two new research papers introduce methods to accelerate AI model inference. The first, Rank-Aware Speculative Sampling (RASS), improves upon existing tree-based speculative sampling techniques for diffusion models by ranking draft candidates and optimizing their selection to minimize discrepancies with the target model's output. RASS demonstrates significant gains, up to 20% faster on CIFAR-10 compared to Diffusion Greedy Rejection Sampling (D-GRS) at matched compute budgets. The second paper presents CAST (Cost-Aware Speculative Trees), which optimizes speculative decoding for large language models by dynamically determining the width of a verification tree based on deployment latency measurements. CAST achieves speedups of up to 43% across various settings and GPU generations, ensuring the target output distribution remains unchanged. AI
IMPACT These techniques could significantly reduce inference latency for diffusion models and large language models, leading to faster and more efficient AI applications.
RANK_REASON Two academic papers introducing novel methods for accelerating AI model inference.
- arXiv
- CAST
- graphics processing unit
- Hugging Face
- large language model
- ScienceCast
- alphaXiv
- CIFAR-10
- COCO2014
- Diffusion Draft Trees
- Diffusion Greedy Rejection Sampling
- FFHQ
- Rank-Aware Speculative Sampling
- Stable Diffusion 3.5
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →