New research tackles KV cache compression for efficient LLM inference · 10 sources tracked
ByPulseAugur Editorial·[24 sources]·
Multiple research papers released on arXiv propose novel methods for optimizing the key-value (KV) cache in large language models to improve inference efficiency and reduce memory usage. Techniques like EchoPress and PatchKV aim to prune or compensate for KV cache compression without significant accuracy loss, while others such as ResidualKV and KV-Kaizen explore adaptive compression strategies. Additionally, methods like ATTUNER and HeadGuard focus on query-side adaptation and selective head protection for specific applications like LLM agents and vision-language models, respectively.
AI
IMPACT
These advancements in KV cache optimization could significantly reduce memory and computational costs for long-context inference, potentially accelerating the deployment of more capable LLMs in resource-constrained environments.
RANK_REASON
Multiple research papers published on arXiv introducing new methods for KV cache compression and optimization in LLMs.
arXiv:2610.00412v1 Announce Type: new Abstract: KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieve…
arXiv:2609.39329v1 Announce Type: new Abstract: Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade …
arXiv:2609.39185v1 Announce Type: new Abstract: Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed bac…
arXiv:2609.32259v2 Announce Type: replace Abstract: Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing th…
arXiv cs.AI
TIER_1English(EN)·Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo Perez, Yiannis Douratsos, Sen He, Zijian Zhou, Fei Zhang, Zhaochong An, Juan-Manuel Perez-Rua, Chen Change Loy, Tao Xiang·
arXiv:2609.32540v2 Announce Type: replace-cross Abstract: Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are alrea…
arXiv cs.AI
TIER_1English(EN)·Jitai Hao, Qiang Huang, Yaowei Wang, Min Zhang, Jun Yu·
arXiv:2602.08005v2 Announce Type: replace-cross Abstract: Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computation grows quadratically. Existing approaches typically address one at the expense …
arXiv:2609.35800v1 Announce Type: new Abstract: Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed h…
arXiv:2609.36835v1 Announce Type: new Abstract: Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. …
arXiv cs.AI
TIER_1English(EN)·Joao Monteiro, Louis B\'ethune, Anastasiia Filippova, Sonia Laguna, David Grangier, Marco Cuturi·
arXiv:2609.37988v1 Announce Type: new Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost …
arXiv cs.AI
TIER_1English(EN)·Xingyu Zhu (Luke), Pu (Luke), Yi, Ziheng Cheng, Ang Lv, Jing Liu, Lexing Ying, Yiyuan Ma, Xin Dong·
arXiv:2609.36322v1 Announce Type: cross Abstract: Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new position…
arXiv cs.CL
TIER_1English(EN)·Xinghao Chen, Junnan Dong, Cai Ke, Chak Tou Leong, Haocheng Sun, Keyu Chen, Siyu An, Ruizhi Qiao, Xing Sun, Wenjie Li, Xiaoyu Shen·
arXiv:2609.36722v1 Announce Type: new Abstract: Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching…
arXiv cs.CL
TIER_1English(EN)·Zunhai Su, Yuxuan Sun, Jianchao Tan, Tao Zhang, Ruihan Hu, Yuchen Xie, Xunliang Cai, Ngai Wong·
arXiv:2609.36760v1 Announce Type: cross Abstract: Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we es…
arXiv cs.LG
TIER_1English(EN)·Jiale Chen, Vage Egiazarian, Eldar Kurti\'c, Torsten Hoefler, Dan Alistarh·
arXiv:2609.38121v1 Announce Type: new Abstract: KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which co…
arXiv:2609.30884v1 Announce Type: new Abstract: Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current mo…
arXiv:2609.30854v1 Announce Type: cross Abstract: Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF…
arXiv:2609.30738v1 Announce Type: cross Abstract: KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispers…
arXiv cs.LG
TIER_1English(EN)·Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona·
arXiv:2609.31415v1 Announce Type: new Abstract: Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fa…
arXiv cs.AI
TIER_1English(EN)·Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou, Teng Wang, Xuehai Zhou·
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-…
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundan…
Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position re…
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundan…
Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less…
<p>A crafted TCP packet to Mooncake's transfer data port is enough to read and write arbitrary process memory — no authentication required.</p> <p>CVE-2026-103764 (CVSS 9.8) is an untrusted pointer dereference in ServerSession::readHeader. The readHeader function trusts attacker-…