Researchers have developed DynaNDE, a dynamic scheduling framework designed to accelerate batched Mixture-of-Experts (MoE) model inference on neural processing unit (NPU)-based systems. This framework leverages near-data processing (NDP) to reduce data movement overhead, a significant bottleneck for MoE models. DynaNDE incorporates an analytical performance model to optimize expert scheduling across NPUs and NDPs, considering hardware heterogeneity and concurrency. Experimental results indicate that DynaNDE can achieve substantial throughput improvements, with average speedups of 2.6x for prefill and 2.2x for decoding stages compared to existing state-of-the-art methods. AI
IMPACT Optimizes MoE inference efficiency, potentially reducing hardware costs and latency for large language models.
RANK_REASON This is a research paper detailing a new framework for optimizing AI model inference. [lever_c_demoted from research: ic=1 ai=1.0]
- AI accelerator
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- DynaNDE
- Gotit.pub
- Hugging Face
- Influence Flower
- mixture of experts
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →