A new research paper explores optimal placement strategies for KV caches across different memory tiers (GPU HBM, CPU DRAM, SSD) to manage scarce GPU memory. The study, conducted using a discrete event simulator, found that tiering memory can support significantly more concurrent sessions and reduce costs, though the placement policy itself had a minimal impact on throughput. Different policies like recency, reuse frequency, and EWMA were evaluated, with reuse frequency performing best for agent and document question answering workloads, while recency was better for chat. AI
IMPACT Optimizing KV cache placement can significantly increase LLM session capacity and reduce operational costs.
RANK_REASON Research paper on optimizing LLM inference infrastructure. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →