English(EN)Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
新研究解决KV缓存压缩问题,以实现高效大语言模型推理 · 追踪10个来源
作者PulseAugur 编辑部·[24 个来源]·
arXiv上发布的多篇研究论文提出了优化大语言模型中键值(KV)缓存的新方法,以提高推理效率并减少内存使用。EchoPress和PatchKV等技术旨在对KV缓存压缩进行剪枝或补偿,而不会显著损失准确性,而ResidualKV和KV-Kaizen等技术则探索自适应压缩策略。此外,ATTUNER和HeadGuard等方法分别侧重于查询端适应和选择性头部保护,以适应LLM智能体和视觉语言模型等特定应用。
AI
arXiv:2610.00412v1 Announce Type: new Abstract: KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieve…
arXiv:2609.39329v1 Announce Type: new Abstract: Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade …
arXiv:2609.39185v1 Announce Type: new Abstract: Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed bac…
arXiv:2609.32259v2 Announce Type: replace Abstract: Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing th…
arXiv cs.AI
TIER_1English(EN)·Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo Perez, Yiannis Douratsos, Sen He, Zijian Zhou, Fei Zhang, Zhaochong An, Juan-Manuel Perez-Rua, Chen Change Loy, Tao Xiang·
arXiv:2609.32540v2 Announce Type: replace-cross Abstract: Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are alrea…
arXiv cs.AI
TIER_1English(EN)·Jitai Hao, Qiang Huang, Yaowei Wang, Min Zhang, Jun Yu·
arXiv:2602.08005v2 Announce Type: replace-cross Abstract: Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computation grows quadratically. Existing approaches typically address one at the expense …
arXiv:2609.35800v1 Announce Type: new Abstract: Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed h…
arXiv:2609.36835v1 Announce Type: new Abstract: Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. …
arXiv cs.AI
TIER_1English(EN)·Joao Monteiro, Louis B\'ethune, Anastasiia Filippova, Sonia Laguna, David Grangier, Marco Cuturi·
arXiv:2609.37988v1 Announce Type: new Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost …
arXiv cs.AI
TIER_1English(EN)·Xingyu Zhu (Luke), Pu (Luke), Yi, Ziheng Cheng, Ang Lv, Jing Liu, Lexing Ying, Yiyuan Ma, Xin Dong·
arXiv:2609.36322v1 Announce Type: cross Abstract: Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new position…
arXiv cs.CL
TIER_1English(EN)·Xinghao Chen, Junnan Dong, Cai Ke, Chak Tou Leong, Haocheng Sun, Keyu Chen, Siyu An, Ruizhi Qiao, Xing Sun, Wenjie Li, Xiaoyu Shen·
arXiv:2609.36722v1 Announce Type: new Abstract: Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching…
arXiv cs.CL
TIER_1English(EN)·Zunhai Su, Yuxuan Sun, Jianchao Tan, Tao Zhang, Ruihan Hu, Yuchen Xie, Xunliang Cai, Ngai Wong·
arXiv:2609.36760v1 Announce Type: cross Abstract: Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we es…
arXiv cs.LG
TIER_1English(EN)·Jiale Chen, Vage Egiazarian, Eldar Kurti\'c, Torsten Hoefler, Dan Alistarh·
arXiv:2609.38121v1 Announce Type: new Abstract: KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which co…
arXiv:2609.30884v1 Announce Type: new Abstract: Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current mo…
arXiv:2609.30854v1 Announce Type: cross Abstract: Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF…
arXiv:2609.30738v1 Announce Type: cross Abstract: KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispers…
arXiv cs.LG
TIER_1English(EN)·Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona·
arXiv:2609.31415v1 Announce Type: new Abstract: Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fa…
arXiv cs.AI
TIER_1English(EN)·Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou, Teng Wang, Xuehai Zhou·
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-…
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundan…
Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position re…
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundan…
Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less…
<p>A crafted TCP packet to Mooncake's transfer data port is enough to read and write arbitrary process memory — no authentication required.</p> <p>CVE-2026-103764 (CVSS 9.8) is an untrusted pointer dereference in ServerSession::readHeader. The readHeader function trusts attacker-…