近期研究探讨了优化大型语言模型(LLM)中KV缓存使用的方法,特别是在长上下文和智能体系统方面。一篇论文提出了一种针对文档编辑后陈旧KV缓存的预算修复策略,发现连续的、与编辑相关的局部窗口非常有效,并且比完全重新预填充快得多。另一项研究介绍了Fathom技术,该技术允许查询动态决定读取多少KV缓存,从而提高解码速度并减少大型模型的注意力误差。第三篇论文研究了KV缓存逐出对自回归生成的影响,分析了发散时间和累积不匹配,发现某些逐出策略会导致更晚的发散和更少的不匹配。最后,关于复制推理共享KV缓存的研究发现了正确性故障和性能边界,在特定场景下显示出显著的加速效果,但也强调了适当验证和局部性条件的重要性。
AI
<p>The decode loop releases MLX's pool of freed buffers every 256 generated<br /> tokens, which is also how often the KV cache grows and drops its previous,<br /> smaller buffers. The check fires only when the token count lands exactly on<br /> a multiple of 256. Speculative deco…
arXiv cs.CL
TIER_1English(EN)·Yi Su, Hong Liu, Guanghua Yu, Jianchen Zhu·
arXiv:2609.19880v1 Announce Type: new Abstract: The KV cache has become a major bottleneck in deploying LLMs, as its memory footprint grows linearly with sequence length and batch size, imposing substantial pressure on both memory capacity and bandwidth. Among various KV cache co…
arXiv:2609.19969v1 Announce Type: new Abstract: The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and lar…
arXiv:2609.17652v1 Announce Type: cross Abstract: When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic that bounds d…
arXiv:2609.17983v1 Announce Type: new Abstract: KV-cache reuse can reduce inference cost in retrieval-augmented generation and agentic systems, but cached contexts may become stale when retrieved knowledge, working memory, or user state is edited. Under causal self-attention, eve…
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capa…
arXiv:2609.16617v1 Announce Type: new Abstract: KV-cache eviction perturbs the conditional token distributions governing autoregressive generation. We investigate how first-divergence timing and subsequent token mismatch determine cumulative disagreement. We derive an exact decom…
KV-cache eviction perturbs the conditional token distributions governing autoregressive generation. We investigate how first-divergence timing and subsequent token mismatch determine cumulative disagreement. We derive an exact decomposition under a specified stepwise maximal coup…
arXiv:2609.15021v1 Announce Type: cross Abstract: Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality. We study two single-GPU 27B vLLM replicas shari…
When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic that bounds decoding. We present Fathom, a key scan in which ea…
dev.to — LLM tag
TIER_1English(EN)·Krishnendu Chatterjee·
<p>I've been working through the inference chapter of a free course I've been studying — <strong><a href="https://ai.studybydoing.in" rel="noopener noreferrer">AI Engineering: Zero to Production</a></strong> — and<br /> the KV-cache lesson finally made something click that I'd be…