A new system called PowerSlider has been developed to optimize LLM serving under fluctuating power constraints, a common issue in AI inference clusters due to demand response requirements. Unlike existing systems that either optimize for static energy or shed fixed priority tiers, PowerSlider exploits the phase asymmetry in LLM workloads. It disaggregates the prefill, think, and answer stages, allowing for per-stage frequency and KV cache control to minimize performance loss when power is capped. This system utilizes a novel Flex SLO contract and an online solver that can re-solve within milliseconds, ensuring sustained goodput and latency-critical tails even under significant power reductions. AI
IMPACT Optimizes LLM serving efficiency under power constraints, potentially reducing operational costs and improving performance during grid emergencies.
RANK_REASON The item describes a novel system and algorithm for optimizing LLM serving, presented in an arXiv paper. [lever_c_demoted from research: ic=1 ai=1.0]
- California Independent System Operator
- demand response
- dynamic voltage and frequency scaling
- Flex SLO
- graphics processing unit
- Karush--Kuhn--Tucker (KKT)
- KV cache
- LLM
- PowerSlider
- SGLang
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →