PulseAugur
中
实时 03:24:12
English(EN) Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

新研究解决KV缓存压缩问题,以实现高效大语言模型推理 · 追踪10个来源

arXiv上发布的多篇研究论文提出了优化大语言模型中键值(KV)缓存的新方法,以提高推理效率并减少内存使用。EchoPress和PatchKV等技术旨在对KV缓存压缩进行剪枝或补偿,而不会显著损失准确性,而ResidualKV和KV-Kaizen等技术则探索自适应压缩策略。此外,ATTUNER和HeadGuard等方法分别侧重于查询端适应和选择性头部保护,以适应LLM智能体和视觉语言模型等特定应用。 AI

影响 这些在KV缓存优化方面的进展可能会显著降低长上下文推理的内存和计算成本,从而有可能加速在资源受限环境中部署更强大的大语言模型。

排序理由 arXiv上发表了多篇研究论文,介绍了大语言模型中KV缓存压缩和优化的新方法。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 24 个来源。 我们如何撰写摘要 →

新研究解决KV缓存压缩问题,以实现高效大语言模型推理 · 追踪10个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
arXiv上发表了多篇研究论文,介绍了大语言模型中KV缓存压缩和优化的新方法。
Source corroboration
24 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
8 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [24]

  1. arXiv cs.LG TIER_1 English(EN) · Jiawei Lin, Saibo Geng, Thomas Bourgeat ·

    EchoPress:通过虚拟上下文重建实现查询无关的KV缓存修剪

    arXiv:2610.00412v1 Announce Type: new Abstract: KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieve…

  2. arXiv cs.LG TIER_1 English(EN) · Chanryeol Lee, Chanhyuk Lee, Yeonwoo Choi, Donggyun Kim, Seunghoon Hong ·

    PatchKV:KV缓存的权重空间补偿

    arXiv:2609.39329v1 Announce Type: new Abstract: Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade …

  3. arXiv cs.LG TIER_1 English(EN) · Snigdha Chandan Khilar ·

    量化递归状态缓存的低差异抖动

    arXiv:2609.39185v1 Announce Type: new Abstract: Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this state in low precision saves memory bandwidth, but every rounding error is fed bac…

  4. arXiv cs.AI TIER_1 English(EN) · Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Murali Annavaram, Sai Praneeth Karimireddy ·

    异构多智能体大语言模型预填充无关的跨家族KV缓存迁移

    arXiv:2609.32259v2 Announce Type: replace Abstract: Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing th…

  5. arXiv cs.AI TIER_1 English(EN) · Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo Perez, Yiannis Douratsos, Sen He, Zijian Zhou, Fei Zhang, Zhaochong An, Juan-Manuel Perez-Rua, Chen Change Loy, Tao Xiang ·

    具有干净锚点的飞行中 KV 缓存,用于加速自回归视频扩散

    arXiv:2609.32540v2 Announce Type: replace-cross Abstract: Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are alrea…

  6. arXiv cs.AI TIER_1 English(EN) · Jitai Hao, Qiang Huang, Yaowei Wang, Min Zhang, Jun Yu ·

    ResidualKV:基于残差的KV缓存压缩,实现高效长上下文推理

    arXiv:2602.08005v2 Announce Type: replace-cross Abstract: Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computation grows quadratically. Existing approaches typically address one at the expense …

  7. arXiv cs.LG TIER_1 English(EN) · Nenad Banfic ·

    HeadGuard:低比特VLM KV缓存量化的选择性头部保护

    arXiv:2609.35800v1 Announce Type: new Abstract: Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed h…

  8. arXiv cs.AI TIER_1 English(EN) · Zheyu Shen, Guanhua Wang, Dezhan Tu, Mengchi Zhang, Yanjia Li, Adnan Aziz, Chunqiang Tang, Ang Li ·

    ARC-KV:用于基于重建的 KV 缓存压缩的摊销锚点搜索

    arXiv:2609.36835v1 Announce Type: new Abstract: Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. …

  9. arXiv cs.AI TIER_1 English(EN) · Joao Monteiro, Louis B\'ethune, Anastasiia Filippova, Sonia Laguna, David Grangier, Marco Cuturi ·

    KV-Kaizen:学习上下文自适应缓存压缩选择

    arXiv:2609.37988v1 Announce Type: new Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost …

  10. arXiv cs.AI TIER_1 English(EN) · Xingyu Zhu (Luke), Pu (Luke), Yi, Ziheng Cheng, Ang Lv, Jing Liu, Lexing Ying, Yiyuan Ma, Xin Dong ·

    周期性弱点:分块 KV 缓存压缩的相位敏感性

    arXiv:2609.36322v1 Announce Type: cross Abstract: Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new position…

  11. arXiv cs.CL TIER_1 English(EN) · Xinghao Chen, Junnan Dong, Cai Ke, Chak Tou Leong, Haocheng Sun, Keyu Chen, Siyu An, Ruizhi Qiao, Xing Sun, Wenjie Li, Xiaoyu Shen ·

    ATTUNER:通过查询端自适应实现无重计算的KV缓存重用

    arXiv:2609.36722v1 Announce Type: new Abstract: Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching…

  12. arXiv cs.CL TIER_1 English(EN) · Zunhai Su, Yuxuan Sun, Jianchao Tan, Tao Zhang, Ruihan Hu, Yuchen Xie, Xunliang Cai, Ngai Wong ·

    QuantMLA:面向低比特MLA KV缓存的函数对齐双路径量化

    arXiv:2609.36760v1 Announce Type: cross Abstract: Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we es…

  13. arXiv cs.LG TIER_1 English(EN) · Jiale Chen, Vage Egiazarian, Eldar Kurti\'c, Torsten Hoefler, Dan Alistarh ·

    WUSH-KV: 具有数据自适应变换的 KV 缓存量化

    arXiv:2609.38121v1 Announce Type: new Abstract: KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which co…

  14. arXiv cs.LG TIER_1 English(EN) · Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen ·

    CacheReforge: 演进式适配器下陈旧 KV 缓存的有界恢复

    arXiv:2609.30884v1 Announce Type: new Abstract: Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current mo…

  15. arXiv cs.LG TIER_1 English(EN) · Tejinder Singh ·

    KV Cache 成为新的内存墙

    arXiv:2609.30854v1 Announce Type: cross Abstract: Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF…

  16. arXiv cs.LG TIER_1 English(EN) · Tianfang Xie, Wei Zhu ·

    超越平均注意力:面向 KV 缓存驱逐的多样性感知、层级式评分

    arXiv:2609.30738v1 Announce Type: cross Abstract: KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $\mu_i+\lambda_1\sigma_i+\lambda_2\mathrm{corr}(i,S)$, adding attention dispers…

  17. arXiv cs.LG TIER_1 English(EN) · Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona ·

    评估KV缓存重用技术的准确性

    arXiv:2609.31415v1 Announce Type: new Abstract: Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fa…

  18. arXiv cs.AI TIER_1 English(EN) · Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou, Teng Wang, Xuehai Zhou ·

    ActKV:通过动作引导的KV缓存管理实现高效LLM代理

    arXiv:2609.31395v1 Announce Type: cross Abstract: Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output qual…

  19. Hugging Face Daily Papers TIER_1 English(EN) ·

    WUSH-KV:具有数据自适应变换的KV缓存量化

    KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-…

  20. Hugging Face Daily Papers TIER_1 English(EN) ·

    异构多智能体大语言模型预填充无关的跨家族KV缓存迁移

    Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundan…

  21. Hugging Face Daily Papers TIER_1 English(EN) ·

    周期性弱点:分块 KV 缓存压缩的相位敏感性

    Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position re…

  22. arXiv cs.MA (Multiagent) TIER_1 English(EN) · Murali Annavaram ·

    异构多智能体大语言模型预填充无关的跨家族KV缓存迁移

    Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundan…

  23. Hugging Face Daily Papers TIER_1 English(EN) ·

    具有清晰锚点的飞行中 KV 缓存,用于更快的自回归视频扩散

    Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, each through a short sequence of denoising stages. To memorize chunks that are already generated, previous methods reconstruct a clean or less…

  24. dev.to — LLM tag TIER_1 English(EN) · threataft ·

    月饼大规模披露 — KV缓存传输引擎中存在 CVSS 9.8 任意内存读/写漏洞

    <p>A crafted TCP packet to Mooncake's transfer data port is enough to read and write arbitrary process memory — no authentication required.</p> <p>CVE-2026-103764 (CVSS 9.8) is an untrusted pointer dereference in ServerSession::readHeader. The readHeader function trusts attacker-…