Researchers have developed a new method to attack the serving frameworks of large language models (LLMs), rather than the models themselves. This "Fill and Squeeze" strategy targets the scheduler's state transitions by exhausting the KV cache and forcing repetitive preemption. The attack has demonstrated significant degradation in Time To First Token (TTFT) and Time Per Output Token on the vLLM framework, proving effective in a practical black-box setting with lower costs than previous methods. AI
IMPACT This research highlights a new class of vulnerabilities in LLM serving infrastructure, potentially impacting deployment security and cost-efficiency.
RANK_REASON Academic paper detailing a new attack methodology on LLM serving frameworks. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →