PulseAugur
中
实时 02:26:51

新研究优化LLM的KV缓存使用,提高效率和准确性

近期研究探讨了优化大型语言模型(LLM)中KV缓存使用的方法,特别是在长上下文和智能体系统方面。一篇论文提出了一种针对文档编辑后陈旧KV缓存的预算修复策略,发现连续的、与编辑相关的局部窗口非常有效,并且比完全重新预填充快得多。另一项研究介绍了Fathom技术,该技术允许查询动态决定读取多少KV缓存,从而提高解码速度并减少大型模型的注意力误差。第三篇论文研究了KV缓存逐出对自回归生成的影响,分析了发散时间和累积不匹配,发现某些逐出策略会导致更晚的发散和更少的不匹配。最后,关于复制推理共享KV缓存的研究发现了正确性故障和性能边界,在特定场景下显示出显著的加速效果,但也强调了适当验证和局部性条件的重要性。 AI

影响 这些KV缓存管理方面的进步可以显著降低推理成本,并提高LLM的性能,尤其是在需要长上下文窗口或智能体行为的应用中。

排序理由 arXiv上发表了多篇研究论文,详细介绍了与LLM中KV缓存优化相关的新方法和分析。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 11 个来源。 我们如何撰写摘要 →

新研究优化LLM的KV缓存使用,提高效率和准确性

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
arXiv上发表了多篇研究论文,详细介绍了与LLM中KV缓存优化相关的新方法和分析。
Source corroboration
11 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
23 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+6 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [11]

  1. Ollama — Releases TIER_1 English(EN) · jessegross ·

    v0.34.2-rc2: mlxrunner: 在推测解码期间释放了已释放的 KV 缓冲区

    <p>The decode loop releases MLX's pool of freed buffers every 256 generated<br /> tokens, which is also how often the KV cache grows and drops its previous,<br /> smaller buffers. The check fires only when the token count lands exactly on<br /> a multiple of 256. Speculative deco…

  2. arXiv cs.CL TIER_1 English(EN) · Yi Su, Hong Liu, Guanghua Yu, Jianchen Zhu ·

    D-Quant:用于 KV 缓存量化的可漂移熵编码

    arXiv:2609.19880v1 Announce Type: new Abstract: The KV cache has become a major bottleneck in deploying LLMs, as its memory footprint grows linearly with sequence length and batch size, imposing substantial pressure on both memory capacity and bandwidth. Among various KV cache co…

  3. arXiv cs.CL TIER_1 English(EN) · DeepSeek-AI, :, Anyi Xu, B. Li, Bangcai Lin, Bing Xue, BingCheng Xian, Bingzheng Xu, Bochao Wu, Bowei Zhang, Boyi Deng, C. C. Yu, Chao Jin, Chaofan Lin, Chen Dong, Chenbing Wang, Chenfan Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyuan Zhang, C… ·

    DeepSeek-V4.1-Flash:突破KV缓存压缩的极限

    arXiv:2609.19969v1 Announce Type: new Abstract: The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and lar…

  4. arXiv cs.CL TIER_1 English(EN) · Vivek Kalyanarangan ·

    Fathom:稀疏解码通过卸载的KV缓存的每查询读取深度

    arXiv:2609.17652v1 Announce Type: cross Abstract: When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic that bounds d…

  5. arXiv cs.AI TIER_1 English(EN) · Mingyang Mao, Wyatt Mackey, Xiaomin Lin ·

    连续性而非重要性:文档编辑后陈旧 KV 缓存的预算修复

    arXiv:2609.17983v1 Announce Type: new Abstract: KV-cache reuse can reduce inference cost in retrieval-augmented generation and agentic systems, but cached contexts may become stale when retrieved knowledge, working memory, or user state is edited. Under causal self-attention, eve…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    DeepSeek-V4.1-Flash:突破KV缓存压缩的极限

    The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capa…

  7. arXiv cs.LG TIER_1 English(EN) · Xinyue Luo, Fei Yu ·

    KV-Cache 驱逐下的发散时机与累积分歧

    arXiv:2609.16617v1 Announce Type: new Abstract: KV-cache eviction perturbs the conditional token distributions governing autoregressive generation. We investigate how first-divergence timing and subsequent token mismatch determine cumulative disagreement. We derive an exact decom…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    KV-Cache 驱逐下的分歧时间和累积不一致

    KV-cache eviction perturbs the conditional token distributions governing autoregressive generation. We investigate how first-divergence timing and subsequent token mismatch determine cumulative disagreement. We derive an exact decomposition under a specified stepwise maximal coup…

  9. arXiv cs.LG TIER_1 English(EN) · Frank Li ·

    复制27B推理的共享KV缓存:正确性失败与性能边界

    arXiv:2609.15021v1 Announce Type: cross Abstract: Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality. We study two single-GPU 27B vLLM replicas shari…

  10. Hugging Face Daily Papers TIER_1 English(EN) ·

    Fathom:稀疏解码在卸载KV缓存上的每查询读取深度

    When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic that bounds decoding. We present Fathom, a key scan in which ea…

  11. dev.to — LLM tag TIER_1 English(EN) · Krishnendu Chatterjee ·

    LLM推理中的KV缓存是什么?(以及为什么是它,而不是权重,限制了你的吞吐量)

    <p>I've been working through the inference chapter of a free course I've been studying — <strong><a href="https://ai.studybydoing.in" rel="noopener noreferrer">AI Engineering: Zero to Production</a></strong> — and<br /> the KV-cache lesson finally made something click that I'd be…