Researchers have formalized the problem of KV cache eviction in large language models, moving beyond heuristic-based methods. By framing KV eviction probabilistically, the problem is reduced to expectation estimation, which can be approximated through sampling. This approach also enables correction for evicted entries during decoding, a previously overlooked issue. The proposed probabilistic method, coupled with decode-time correction, demonstrates greater robustness across various tasks and achieves competitive performance. AI
IMPACT Formalizes KV cache eviction, potentially improving LLM inference efficiency and throughput.
RANK_REASON The cluster contains a research paper published on arXiv detailing a new theoretical approach to a technical problem in LLM inference. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- KV cache
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →