PulseAugur
实时 23:09:03
English(EN) BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

新研究致力于 LLM KV 缓存压缩,实现高效长上下文推理 · 跟踪 5 个来源

arXiv 上发布的多个研究论文提出了压缩大型语言模型键值(KV)缓存的新颖方法,以解决长上下文推理期间的内存瓶颈。BeaconKVSGD-KVVestigeKVRandom Attention 等技术旨在通过智能选择或压缩 KV 缓存数据来减少内存使用并提高吞吐量。这些方法包括使用信标查询和摘要引导的诊断、随机驱逐以及利用残余信号,所有这些方法都力求在显著降低内存需求的同时保持准确性。 AI

影响 KV 缓存压缩方面的这些进展可以显著提高大型语言模型的效率和可扩展性,从而实现更复杂的推理和更长的上下文窗口。

排序理由 arXiv 上发布了多篇研究论文,介绍了 LLM KV 缓存压缩的新方法。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 15 个来源。 我们如何撰写摘要 →

新研究致力于 LLM KV 缓存压缩,实现高效长上下文推理 · 跟踪 5 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
arXiv 上发布了多篇研究论文,介绍了 LLM KV 缓存压缩的新方法。
Source corroboration
15 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
9 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [15]

  1. arXiv cs.CL TIER_1 English(EN) · Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi ·

    BeaconKV:由 Beacon 查询引导的键值缓存压缩,用于高效的大型推理模型推理

    arXiv:2609.04971v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, o…

  2. arXiv cs.CL TIER_1 English(EN) · Seifeldin Abdellatif ·

    通过低秩注意力适应实现量化 KV 缓存的质量恢复

    arXiv:2609.04263v1 Announce Type: cross Abstract: Low-bit key--value (KV) caches reduce the memory required for autoregressive decoding, but the resulting quality loss depends on the model and quantizer. We keep the quantizer fixed and distill the floating-cache model's behavior …

  3. arXiv cs.CL TIER_1 English(EN) · WenJie Fan ·

    VestigeKV:NoPE-MLA KV缓存自带逐出信号,位于残余分支

    arXiv:2609.03949v1 Announce Type: cross Abstract: The problem. A long-lived KV cache must be compressed before the queries that will read it exist; selection by observed attention (H2O, SnapKV) collapses there (0.00-0.33 needle retrieval on a NoPE MLA model), because a token's im…

  4. arXiv cs.CL TIER_1 English(EN) · Zeyu Liu, Woomin Song, Xuandi Fu, Sai Muralidhar Jayanthi, Vivek Govindan, Aram Galstyan, Sravan Babu Bodapati, Srikanth Ronanki ·

    SGD-KV:摘要引导的KV缓存压缩

    arXiv:2609.03235v1 Announce Type: new Abstract: Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooki…

  5. arXiv cs.CL TIER_1 English(EN) · Heng Wang, Jielin Qiu, Wenting Zhao, Cheng Qian, Liangwei Yang, Jiawei Han, Heng Ji, Silvio Savarese, Shelby Heinecke, Huan Wang ·

    随机注意力:重新思考KV缓存逐出以实现高效推理

    arXiv:2609.03430v1 Announce Type: new Abstract: Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score ea…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    BeaconKV:由Beacon查询引导的键值缓存压缩,实现高效大型推理模型推理

    BeaconKV improves memory efficiency for long reasoning traces by using compact beacon queries to predict which past key-value pairs will be revisited, reducing cache size without sacrificing accuracy.

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    SGD-KV:摘要引导的KV缓存压缩

    Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooking the distinct functional roles of different at…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    随机注意力:重新思考 KV 缓存逐出以实现高效推理

    Random eviction of reasoning tokens matches selective KV cache compression because reasoning traces are self-protecting through redundancy, making scoring unnecessary once prompts are preserved.

  9. arXiv cs.AI TIER_1 English(EN) · Xingyu Qu, Siyuan Lu, Zhiyu Chen, Sheng Wang, Tao Lin ·

    CacheBridge:高效跨模型KV缓存传输

    arXiv:2609.00891v1 Announce Type: new Abstract: Sharing context between LLMs in a multi-model system requires the receiving model to prefill the shared prefix because KV caches are model-specific. Recent closed-form cross-model KV transfer, hereafter Full-Head Mapping, avoids thi…

  10. arXiv cs.AI TIER_1 English(EN) · Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev, Vyacheslav Zhdanovskiy, Yegor Yershov ·

    KV Cache Offloading for Context-Intensive Tasks

    arXiv:2604.08426v5 Announce Type: replace-cross Abstract: With the growing demand for long-context LLMs across a wide range of applications, the key-value (KV) cache has become a critical bottleneck for both latency and memory usage. Recently, KV-cache offloading has emerged as a…

  11. arXiv cs.CL TIER_1 English(EN) · Renato Geh, Alex Chen, Daniel Israel, Aditya Grover, Guy Van den Broeck ·

    KV 缓存逐出的概率解释

    arXiv:2608.28293v1 Announce Type: new Abstract: The premise and promise of KV (cache) eviction is simple: higher throughput can be achieved by evicting some entries from the KV cache, at a negligible cost to quality. This holds empirically for many existing methods, though most r…

  12. Towards AI TIER_1 English(EN) · Naveen ·

    KV Cache:主导 GPU 内存的 AI 不可见数据库

    <div class="medium-feed-item"><p class="medium-feed-image"><a href="https://pub.towardsai.net/the-kv-cache-ais-unseen-database-dominating-gpu-memory-8b6965957e47?source=rss----98111c9905da---4"><img src="https://cdn-images-1.medium.com/max/1376/1*rBT6aHjQjuz0ZpBqRqIO1g.png" width…

  13. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    小型模型推理如何通过KV缓存降低延迟和能耗

    <p>The latency and energy bottleneck in small-model inference often lies not in compute but in memory access and VRAM management. A well-designed KV Cache tiering and disaggregated storage architecture can improve both time-to-first-token (TTFT) and energy per unit of throughput.…

  14. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    KV缓存空间局部性优化:从分页到分层存储

    <p>Spatial locality optimization for KV Cache is evolving from in-HBM paging management to a tiered architecture spanning multiple storage layers. Measured on the Mingxin FX100 under a 480B production deployment with long-context cold-recovery workloads, tiered acceleration impro…

  15. r/LocalLLaMA TIER_1 English(EN) · /u/giveen ·

    Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen · Pull Request #357 · TheTom/llama-cpp-turboquant

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1w8jflp/block_kv_cache_streaming_bound_vram_at_long/"> <img alt="Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen · Pull Request #357 · TheTom/llama-cpp-turboquant" …