PulseAugur
EN
LIVE 04:05:12

New research tackles KV cache compression for efficient LLM inference · 10 sources tracked

Multiple research papers released on arXiv propose novel methods for optimizing the key-value (KV) cache in large language models to improve inference efficiency and reduce memory usage. Techniques like EchoPress and PatchKV aim to prune or compensate for KV cache compression without significant accuracy loss, while others such as ResidualKV and KV-Kaizen explore adaptive compression strategies. Additionally, methods like ATTUNER and HeadGuard focus on query-side adaptation and selective head protection for specific applications like LLM agents and vision-language models, respectively. AI

IMPACT These advancements in KV cache optimization could significantly reduce memory and computational costs for long-context inference, potentially accelerating the deployment of more capable LLMs in resource-constrained environments.

RANK_REASON Multiple research papers published on arXiv introducing new methods for KV cache compression and optimization in LLMs.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 24 sources. How we write summaries →

New research tackles KV cache compression for efficient LLM inference · 10 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Multiple research papers published on arXiv introducing new methods for KV cache compression and optimization in LLMs.
Source corroboration
24 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
8 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [24]

  1. arXiv cs.LG TIER_1 English(EN) · Jiawei Lin, Saibo Geng, Thomas Bourgeat ·

    EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction

    arXiv:2610.00412v1 Announce Type: new Abstract: KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieve…

  2. arXiv cs.LG TIER_1 English(EN) · Chanryeol Lee, Chanhyuk Lee, Yeonwoo Choi, Donggyun Kim, Seunghoon Hong ·

    PatchKV: Weight-Space Compensation of KV Cache

    arXiv:2609.39329v1 Announce Type: new Abstract: Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade …

  3. arXiv cs.LG TIER_1 English(EN) · Snigdha Chandan Khilar ·

    Low-Discrepancy Dither for Quantized Recurrent State Caches

    arXiv:2609.39185v1 Announce Type: new Abstract: Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed bac…

  4. arXiv cs.AI TIER_1 English(EN) · Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Murali Annavaram, Sai Praneeth Karimireddy ·

    Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

    arXiv:2609.32259v2 Announce Type: replace Abstract: Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing th…

  5. arXiv cs.AI TIER_1 English(EN) · Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo Perez, Yiannis Douratsos, Sen He, Zijian Zhou, Fei Zhang, Zhaochong An, Juan-Manuel Perez-Rua, Chen Change Loy, Tao Xiang ·

    In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

    arXiv:2609.32540v2 Announce Type: replace-cross Abstract: Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are alrea…

  6. arXiv cs.AI TIER_1 English(EN) · Jitai Hao, Qiang Huang, Yaowei Wang, Min Zhang, Jun Yu ·

    ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference

    arXiv:2602.08005v2 Announce Type: replace-cross Abstract: Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computation grows quadratically. Existing approaches typically address one at the expense …

  7. arXiv cs.LG TIER_1 English(EN) · Nenad Banfic ·

    HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization

    arXiv:2609.35800v1 Announce Type: new Abstract: Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed h…

  8. arXiv cs.AI TIER_1 English(EN) · Zheyu Shen, Guanhua Wang, Dezhan Tu, Mengchi Zhang, Yanjia Li, Adnan Aziz, Chunqiang Tang, Ang Li ·

    ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction

    arXiv:2609.36835v1 Announce Type: new Abstract: Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. …

  9. arXiv cs.AI TIER_1 English(EN) · Joao Monteiro, Louis B\'ethune, Anastasiia Filippova, Sonia Laguna, David Grangier, Marco Cuturi ·

    KV-Kaizen: Learning Context-Adaptive Cache Compression Choices

    arXiv:2609.37988v1 Announce Type: new Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost …

  10. arXiv cs.AI TIER_1 English(EN) · Xingyu Zhu (Luke), Pu (Luke), Yi, Ziheng Cheng, Ang Lv, Jing Liu, Lexing Ying, Yiyuan Ma, Xin Dong ·

    Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

    arXiv:2609.36322v1 Announce Type: cross Abstract: Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new position…

  11. arXiv cs.CL TIER_1 English(EN) · Xinghao Chen, Junnan Dong, Cai Ke, Chak Tou Leong, Haocheng Sun, Keyu Chen, Siyu An, Ruizhi Qiao, Xing Sun, Wenjie Li, Xiaoyu Shen ·

    ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation

    arXiv:2609.36722v1 Announce Type: new Abstract: Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching…

  12. arXiv cs.CL TIER_1 English(EN) · Zunhai Su, Yuxuan Sun, Jianchao Tan, Tao Zhang, Ruihan Hu, Yuchen Xie, Xunliang Cai, Ngai Wong ·

    QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching

    arXiv:2609.36760v1 Announce Type: cross Abstract: Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we es…

  13. arXiv cs.LG TIER_1 English(EN) · Jiale Chen, Vage Egiazarian, Eldar Kurti\'c, Torsten Hoefler, Dan Alistarh ·

    WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

    arXiv:2609.38121v1 Announce Type: new Abstract: KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which co…

  14. arXiv cs.LG TIER_1 English(EN) · Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen ·

    CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters

    arXiv:2609.30884v1 Announce Type: new Abstract: Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current mo…

  15. arXiv cs.LG TIER_1 English(EN) · Tejinder Singh ·

    The KV Cache Is the New Memory Wall

    arXiv:2609.30854v1 Announce Type: cross Abstract: Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF…

  16. arXiv cs.LG TIER_1 English(EN) · Tianfang Xie, Wei Zhu ·

    Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction

    arXiv:2609.30738v1 Announce Type: cross Abstract: KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispers…

  17. arXiv cs.LG TIER_1 English(EN) · Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona ·

    Evaluating the accuracy of KV cache reuse techniques

    arXiv:2609.31415v1 Announce Type: new Abstract: Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fa…

  18. arXiv cs.AI TIER_1 English(EN) · Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou, Teng Wang, Xuehai Zhou ·

    ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

    arXiv:2609.31395v1 Announce Type: cross Abstract: Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output qual…

  19. Hugging Face Daily Papers TIER_1 English(EN) ·

    WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

    KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-…

  20. Hugging Face Daily Papers TIER_1 English(EN) ·

    Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

    Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundan…

  21. Hugging Face Daily Papers TIER_1 English(EN) ·

    Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

    Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position re…

  22. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Murali Annavaram ·

    Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

    Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundan…

  23. Hugging Face Daily Papers TIER_1 English(EN) ·

    In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion

    Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less…

  24. dev.to — LLM tag TIER_1 English(EN) · threataft ·

    Mooncake Mass Disclosure — CVSS 9.8 Arbitrary Memory Read/Write in KV Cache Transfer Engine

    <p>A crafted TCP packet to Mooncake's transfer data port is enough to read and write arbitrary process memory — no authentication required.</p> <p>CVE-2026-103764 (CVSS 9.8) is an untrusted pointer dereference in ServerSession::readHeader. The readHeader function trusts attacker-…