PulseAugur
实时 08:48:14
English(EN) REAL: REtrieval-reAsoning and Logic-constructed Attention Behaviors for Long-Context KV Cache Compression

新方法旨在提高 LLM KV 缓存压缩效率

研究人员正在开发先进的技术来压缩大型语言模型 (LLM) 的键值 (KV) 缓存,这是推理过程中内存成本的主要贡献者。JoLTFlashJoLT 等新方法利用张量分解和残差分配,在性能损失极小的情况下实现了显著压缩,如在 Mistral-7B-v0.3Llama 2 13B 等模型上所证明的。其他研究探讨了查询可见性如何影响压缩方法的排名,一些方法在事先不知道查询的情况下表现不佳。此外,PM-KVQREAL 等方法旨在处理长上下文窗口和多样化的注意力行为,以减少累积误差并提高长 CoT LLM 的效率。 AI

影响 KV 缓存压缩的进步可以显著降低推理成本和内存需求,从而实现更高效的 LLM 部署,尤其是在长上下文任务中。

排序理由 多篇研究论文详细介绍了压缩 LLM KV 缓存的新颖方法。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 10 个来源。 我们如何撰写摘要 →

新方法旨在提高 LLM KV 缓存压缩效率

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文详细介绍了压缩 LLM KV 缓存的新颖方法。
Source corroboration
10 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
infra, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
50 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+2 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [10]

  1. arXiv cs.AI TIER_1 English(EN) · Donghyun Son, Euntae Choi, Sungjoo Yoo ·

    NSNQuant:KV缓存的无校准低比特向量量化双重归一化方法

    arXiv:2505.18231v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache. Vector Quantization (VQ) is recently adopt…

  2. arXiv cs.AI TIER_1 English(EN) · Daming Luo, Christy Liang, Junyu Xuan ·

    查询可见性如何改变 KV 缓存压缩排名:一项匹配预算审计

    arXiv:2607.11942v1 Announce Type: cross Abstract: KV-cache compression methods are predominantly evaluated with the query appended to the context before compression -- a query-aware protocol. Yet the economic case for a compressed KV cache is reuse: compress a document once, answ…

  3. arXiv cs.CL TIER_1 English(EN) · Rahul Krishnan, Volker Schulz ·

    A JoLT for the KV Cache: 联合Tucker和JL-Residual分配实现LLM的近乎无损KV Cache压缩

    arXiv:2607.12550v1 Announce Type: cross Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference. It grows with batch size, context length, and depth, and at long context it, rather than the model weights, sets the ceiling on throughput. Two…

  4. arXiv cs.CL TIER_1 English(EN) · Volker Schulz ·

    A JoLT for the KV Cache: 联合Tucker和JL-Residual分配实现LLM近乎无损的KV Cache压缩

    The key-value (KV) cache has become the dominant memory cost of transformer inference. It grows with batch size, context length, and depth, and at long context it, rather than the model weights, sets the ceiling on throughput. Two families of methods reduce it. Low-rank methods f…

  5. arXiv cs.AI TIER_1 English(EN) · Paolo D'Alberto, Ashish Siarasao, Elliott Delaye, Rajeev Patwari ·

    KV缓存压缩的消融、统计推断和验证

    arXiv:2607.09683v1 Announce Type: cross Abstract: This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with Beta Lloyd-Max and QJL, through a statistical validation methodology that separat…

  6. arXiv cs.CL TIER_1 English(EN) · Tengxuan Liu, Shiyao Li, Jiayi Yang, Tianchen Zhao, Feng Zhou, Xiaohui Song, Guohao Dai, Shengen Yan, Huazhong Yang, Yu Wang ·

    PM-KVQ:面向长上下文CoT大模型的渐进式混合精度KV缓存量化

    arXiv:2505.18610v2 Announce Type: replace Abstract: Recently, significant progress has been made in developing reasoning-capable Large Language Models (LLMs) through long Chain-of-Thought (CoT) techniques. However, this long-CoT reasoning process imposes substantial memory overhe…

  7. arXiv cs.AI TIER_1 English(EN) · Mengjie Li, Yuan Feng, Xike Xie, William J. Song ·

    REAL:用于长上下文KV缓存压缩的检索-推理和逻辑构建注意力行为

    arXiv:2508.15806v2 Announce Type: replace-cross Abstract: The growing sequence length of large language models poses significant challenges for key-value (KV) caches. Existing state-of-the-art cache eviction methods primarily analyze the inference behavior of attention heads in s…

  8. Towards AI TIER_1 English(EN) · Mohit Sewak, Ph.D. ·

    超越KV缓存:接下来是什么

    <h4>An outlook on how input-dependent step-size updates will shape the next generation of stream-processing AI.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*gaLdwPNUza_RP-Gn" /></figure><p><em>Cover infographic visualizing the evolutionary leap from hea…

  9. Medium — MLOps tag TIER_1 English(EN) · Jagadish Mukku ·

    量化LLM KV缓存内存扩展的NVMe存储需求

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/@jagadish.mukku/quantifying-nvme-storage-requirements-for-llm-kv-cache-memory-extension-6c4e14d5f102?source=rss------mlops-5"><img src="https://cdn-images-1.medium.com/max/1022/1*iu4fxGbxvZ2DVO…

  10. dev.to — LLM tag TIER_1 English(EN) · AITinkerer ·

    KV Cache 实际节省多少成本?中国大模型缓存定价实测

    <p>Most developers know KV cache reduces costs. Few have actually modeled how much, at current pricing, cache hit rate moves the needle on their bill.</p> <p>Take Qwen3.7 Plus on RouteAI: standard input at <strong>$0.24/M tokens</strong>, cache read at <strong>$0.048/M</strong> —…