PulseAugur
中
实时 23:35:29
English(EN) DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation

新研究致力于 LLM KV 缓存优化以提高效率 · 已追踪 10 个来源

近期研究论文引入了优化大型语言模型中 KV 缓存管理的新技术,解决了内存瓶颈并提高了推理效率。vToken、GCache、LinearKV、KVDiagnosis、SemPIC、SPECTRA、CommitKV、RippleKV、CoinRAG 和 GraceKV 等方法提出了各种策略,包括令牌级虚拟化、全局影响缓存、与位置无关的缓存、诊断基准测试、语义缓存、谱变换编码、生命周期感知压缩和全局资源分配。这些方法旨在减少内存使用量、加速处理并增强 LLM 的性能,特别是在长上下文和检索增强生成任务方面。 AI

影响 KV 缓存管理方面的这些进步对于实现更大的上下文窗口和更高效的 LLM 推理至关重要,直接影响复杂代理和检索增强生成任务的可行性。

排序理由 多篇在 arXiv 上发表的研究论文介绍了 LLM 中 KV 缓存优化的新方法。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 55 个来源。 我们如何撰写摘要 →

新研究致力于 LLM KV 缓存优化以提高效率 · 已追踪 10 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇在 arXiv 上发表的研究论文介绍了 LLM 中 KV 缓存优化的新方法。
Source corroboration
55 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
infra, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
71 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [55]

  1. arXiv cs.CL TIER_1 English(EN) · Hannah Laus, Claudio Mayrink Verdun, Hao Wang, Flavio du Pin Calmon, Felix Krahmer ·

    KV Cache 压缩的变换编码视角

    arXiv:2608.14191v1 Announce Type: cross Abstract: The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference. Existing quantization methods address this bottleneck by representing the KV cache uniformly with lower-preci…

  2. arXiv cs.AI TIER_1 English(EN) · Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue, Cheng Jin, Weizhong Zhang ·

    从局部失配到全球影响:优化缓存重用策略以实现高效扩散

    arXiv:2608.13043v1 Announce Type: new Abstract: Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration has emerged as a promising solution, existing policies rely on local similarity …

  3. arXiv cs.AI TIER_1 English(EN) · Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li ·

    vToken: 用于可回收 KV 缓存的令牌级虚拟化

    arXiv:2608.13263v1 Announce Type: new Abstract: Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragmentation, but recent KV eviction al…

  4. arXiv cs.AI TIER_1 English(EN) · Yirui Liu, Ruoling Qi, Longwen Wang, Xuaner Wu, Jian Chen, Yuxin Jin, Jiawei Shao, Xuelong Li ·

    LinearKV:混合 LLM 中位置无关缓存的单一缓存状态足以实现缓存

    arXiv:2608.11231v1 Announce Type: new Abstract: LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache underlies its core operations: matching reusable token…

  5. arXiv cs.AI TIER_1 English(EN) · Chen Qiu, Ziwu Liu, Chao Fei, Guozhong Li, Panos Kalnis ·

    KVDiagnosis:长上下文语言模型KV缓存压缩的诊断基准

    arXiv:2608.09412v1 Announce Type: new Abstract: KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct executions fail nor why. We present KVDiagnosis, a diagnostic dataset and benchmark with three contributions. First, a 25-metho…

  6. arXiv cs.LG TIER_1 English(EN) · Jiamu Zhang, Liang Wu, Kelly Wan, Hanjie Chen, Liangjie Hong ·

    SPECTRA:通过谱变换编码将 KV 缓存推向 2 位悬崖之外

    arXiv:2608.07915v1 Announce Type: new Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Their inference memory is then dominated by the key-value (KV) cache, the stored a…

  7. arXiv cs.LG TIER_1 English(EN) · Weizhong Huang, Jinchao Zhang, Xiawu Zheng ·

    CommitKV:通过提交转换实现生命周期感知的 KV 缓存压缩,用于多轮对话代理

    arXiv:2608.07855v1 Announce Type: new Abstract: Multi-turn Reasoning-and-Acting (ReAct) agents accumulate growing trajectories of reasoning, tool calls, and observations. Their key-value (KV) caches grow accordingly, increasing memory use and attention cost during model inference…

  8. arXiv cs.AI TIER_1 English(EN) · Hui Xie, Peng Xiao, Yutong Deng, Shuoran Dou, Jian Yang, Jinyang Guo ·

    SemPIC: 学习语义位置无关KV缓存

    arXiv:2607.28069v2 Announce Type: replace Abstract: Long-context retrieval and agentic workloads repeatedly reuse the same documents under changing instructions, histories, and document orders. Prefix caching cannot exploit this reuse, while position-independent caching (PIC) rem…

  9. arXiv cs.AI TIER_1 English(EN) · Dongjie Xu, Kai Qian, Julius, Weijie Shi, Yuxuan Sun, Minghua Tang, Fenglei Jin, Hanchi Dong, Jiajie Xu ·

    RippleKV:通过扰动传播实现跨层 KV 缓存分配

    arXiv:2608.08684v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. Existing methods rely on proxies such as layer depth, attention statistics, or representatio…

  10. arXiv cs.LG TIER_1 English(EN) · Haolin Tian, Yuzhe Liu, Tonghan Wang ·

    每个缓存条目都物有所值:KV缓存压缩的全局分配、解析和覆盖范围

    arXiv:2608.07001v1 Announce Type: new Abstract: As large language models (LLMs) process increasingly long contexts, KV cache storage and repeated access have become a major bottleneck. Existing KV cache compression methods rely on predefined, fixed compression rules and are typic…

  11. arXiv cs.AI TIER_1 English(EN) · Gyuwan Kim, Cheoneum Park, Tao Yang ·

    CoinRAG: 长上下文RAG的上下文信息片段KV缓存重用

    arXiv:2608.07458v1 Announce Type: cross Abstract: Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise st…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    OasisKV:通过前瞻性稀疏预取将解码内 KV 缓存扩展到 HBM 之外

    Large language model (LLM) inference serving is increasingly constrained by memory rather than compute. As long-context and long-form reasoning workloads become more prevalent, the key-value (KV) cache dominates both memory footprint and memory traffic during LLM token generation…

  13. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Tao Yang ·

    CoinRAG:长上下文RAG的上下文信息片段KV缓存重用

    Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This pape…

  14. Hugging Face Daily Papers TIER_1 English(EN) ·

    CoinRAG: 长上下文RAG的上下文信息片段KV缓存复用

    CoinRAG improves retrieval-augmented generation efficiency and accuracy by reusing fine-grained semantic nugget caches instead of full chunks.

  15. arXiv cs.AI TIER_1 English(EN) · Samuel Fern\'andez-Mendui\~na, Amir Ziashahabi, Eduardo Pavez, Antonio Ortega, Salman Avestimehr ·

    在查询时花费比特:具有保持注意力变换的KV缓存向量量化

    arXiv:2608.04074v1 Announce Type: cross Abstract: Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is bandwidth-bound. Hence, reducing the cache size can raise both decoding speed and s…

  16. arXiv cs.AI TIER_1 English(EN) · Wonpyo Park, Seung-won Hwang ·

    TaskPress:通过任务引导剪枝实现查询无关的 KV 缓存压缩

    arXiv:2608.03276v1 Announce Type: new Abstract: Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length. While pruning offers mitigation, prevailing methods determine query-specific token importance that cann…

  17. arXiv cs.CL TIER_1 English(EN) · Malik Khalaf, Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki, Assaf Schuster ·

    AnchorKV:Anchor-Residual KV 缓存压缩

    arXiv:2608.02901v1 Announce Type: cross Abstract: The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction methods permanently discard tokens, degrading performance whenever a discarded tok…

  18. arXiv cs.AI TIER_1 English(EN) · Vincent-Daniel Yun, Woosang Lim, Minsoo Cheong, Sunwoo Lee, Murali Annavaram, Sai Praneeth Karimireddy, Sungjoo Yoo ·

    面向INT2 KV缓存量化的输出感知旋转

    arXiv:2608.02691v1 Announce Type: cross Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important. However, existing rotation-based INT2 methods op…

  19. arXiv cs.LG TIER_1 English(EN) · Taekyung Heo, Rasoul Shafipour, Ritchie Zhao, Maximilian Golub, Mohammad Mahdi Kamani, Ritika Borkar, Makesh Tarun Chandran, Pantea Zardoshti, Bita Darvish Rouhani ·

    大型语言模型家族中的跨模型KV缓存迁移:用于预填充重用的闭式线性映射

    arXiv:2608.03893v1 Announce Type: new Abstract: Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch. We propose cross-…

  20. arXiv cs.LG TIER_1 English(EN) · Jihao Xin, Tian Lyu, David Keyes, Hatem Ltaief, Marco Canini ·

    RAP:通过RoPE对齐剪枝实现KV缓存压缩

    arXiv:2602.02599v4 Announce Type: replace Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the memory and compute of the key-value (KV) cache. Structured pruning is a direct way to shrink it: dropping the least useful channels of the W_k, W_v pr…

  21. arXiv cs.CL TIER_1 English(EN) · Jialong Han, You Wu, Kewei Tu ·

    S$^4$R:选择性采样、子空间和稀疏重构用于压缩长上下文 KV 缓存

    arXiv:2608.00528v1 Announce Type: new Abstract: The growth of context window lengths in Large Language Models (LLMs) significantly enhances their long-context capabilities but incurs prohibitive memory costs due to the Key-Value (KV) cache. Although low-rank compression of KV cac…

  22. arXiv cs.CL TIER_1 English(EN) · Changwoo Baek, Seungjun Shin, Kyeongbo Kong ·

    RestoreKV:在激进的与查询无关的KV缓存驱逐下恢复完整的缓存行为

    arXiv:2608.01247v1 Announce Type: new Abstract: Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse under tight budgets. Existing methods primarily improve which original KV pairs are…

  23. arXiv cs.LG TIER_1 English(EN) · Omin Kwon, Doyeon Kim, Jongseok Park, Seung Yul Lee, Ion Stoica, Jae W. Lee ·

    HERALD:通过 CPU-GPU 协同 KV 缓存检索实现高吞吐量块扩散 LLM 服务

    arXiv:2606.21633v2 Announce Type: replace Abstract: The KV cache dominates GPU memory in long-context LLM serving, crowding out batch capacity and leaving GPU compute idle. Offloading the cache to CPU DRAM restores capacity, but the limited PCIe bandwidth forces state-of-the-art …

  24. arXiv cs.CL TIER_1 English(EN) · Yuhang Zhan, Lisi Chen, Shuo Shang ·

    ResKV:重建被忽略的注意力贡献以实现固定预算的KV缓存压缩

    arXiv:2607.29591v1 Announce Type: new Abstract: KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently remove their aggregate contribution to attention. Merging-based alternatives pr…

  25. Hugging Face Daily Papers TIER_1 English(EN) ·

    RestoreKV:在激进的查询无关KV缓存驱逐下恢复全缓存行为

    Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse under tight budgets. Existing methods primarily improve which original KV pairs are retained. We introduce RestoreKV, which complem…

  26. Hugging Face Daily Papers TIER_1 English(EN) ·

    RestoreKV:在激进的查询无关KV缓存驱逐下恢复全缓存行为

    Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse under tight budgets. Existing methods primarily improve which original KV pairs are retained. We introduce RestoreKV, which complem…

  27. arXiv cs.LG TIER_1 English(EN) · Stephen Gould, Anton van den Hengel ·

    穿越未来:反因果惊喜的键值缓存管理

    arXiv:2607.27600v1 Announce Type: new Abstract: Key-value (KV) cache management through compression and eviction strategies has emerged as an important research direction in recent years. Computational demands of large language models (LLMs) and their multi-modal variants during …

  28. arXiv cs.CL TIER_1 English(EN) · Alexander Boesgaard Lorup ·

    KV缓存导致阶段重放分歧:固定前缀精度控制与双向缓存移植

    arXiv:2607.28495v1 Announce Type: cross Abstract: Stage-replay diagnostics reconstruct intermediate token prefixes and treat fresh-prefill continuation as continuation from the decoder state that originally reached the prefix. We audit that assumption at a whole reasoning-stage b…

  29. arXiv cs.LG TIER_1 English(EN) · Tan T. Nguyen, Quan V. Dang ·

    DynaCalKV:通过头分组和自适应秩分配实现键值缓存压缩

    arXiv:2607.24331v1 Announce Type: new Abstract: As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually becomes a significant bottleneck as the context windo…

  30. arXiv cs.LG TIER_1 English(EN) · Chao Fang, Jun Yin, Man Shi, Marian Verhelst ·

    HiKV:用于大语言模型解码的具有硬件加速的分层重要性感知KV缓存

    arXiv:2607.22389v1 Announce Type: cross Abstract: With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. To tackle this challenge, we propose HiKV, a novel algorithm-hardwa…

  31. arXiv cs.AI TIER_1 English(EN) · Jingquan Chen, Jinghua Piao, Jie Feng, Shaogang Hu, Yong Li ·

    MiniCache:使用小型模型接口的可重用程序缓存,用于高效的 LLM 推理

    arXiv:2607.20507v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for program-aided reasoning, agentic decision making, and structured task execution, but these applications often incur high inference cost. We present MiniCache, a reusable program…

  32. arXiv cs.AI TIER_1 English(EN) · Pragaash Ponnusamy, Shivam Sahni, Jue Wang, Tri Dao ·

    SonicSampler:用于 LLM 采样和推测性验证的统一瓦片感知内核

    arXiv:2607.20475v1 Announce Type: new Abstract: Sampling in LLM inference comprises a combinatorial set of logit processing, token selection, and verification operations for speculative decoding. However, existing implementations either accelerate only subsets of this pipeline, r…

  33. arXiv cs.CL TIER_1 English(EN) · Chuheng Du, Junyi Chen, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Chaoyue Niu, Shengzhong Liu, Guihai Chen, Fan Wu ·

    C$^2$KV:用于高效LLM推理的压缩和可组合KV缓存重用

    arXiv:2607.17715v1 Announce Type: new Abstract: Long-context inference is central to modern large language model (LLM) applications such as retrieval-augmented generation and multi-document reasoning. To mitigate the growing inference cost, recent work has explored key-value (KV)…

  34. arXiv cs.AI TIER_1 English(EN) · Yan Song ·

    缓存感知提示压缩:LLM API缓存的双层成本模型

    arXiv:2607.15516v1 Announce Type: cross Abstract: Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent). The compression literature has standardized on query-aware…

  35. arXiv cs.CL TIER_1 English(EN) · Shahrzad Esmat, Dhawal Shah, Ali Jannesari ·

    VarRate:面向长上下文大语言模型的无训练可变速率KV缓存压缩

    arXiv:2607.15498v1 Announce Type: new Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance…

  36. dev.to — MCP tag TIER_1 English(EN) · Kasi Yaswanth ·

    使用 MCP 缓存优化延迟

    <p>I recently spent a frustrating afternoon debugging a support bot built on top of LangGraph, where the bot would occasionally take an inordinate amount of time to respond to user queries. The bot's workflow involved multiple tool calls to external services, such as entity disam…

  37. Towards AI TIER_1 English(EN) · Jordan Carson ·

    安全带,第二部分:缓存经济学

    <h4>It was never <em>only</em> about tokens. It’s all about the cache.</h4><blockquote>Read the article for free <a href="https://medium.com/@jordancarson/af9e6b92139e?source=friends_link&amp;sk=d1099fdd2f99ae869b92f3aefce68aab">here</a>.</blockquote><figure><img alt="" src="http…

  38. Towards AI TIER_1 English(EN) · Deepanshu Gupta ·

    上下文正成为基础设施:从KV缓存到矛盾感知RAG

    <h4>The next generation of LLM systems will be shaped not only by their models, but by how well they manage context.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*LDZLP4ysvvqcmhlnz6alpg.png" /></figure><p>Imagine a production AI assistant that retrieves …

  39. dev.to — LLM tag TIER_1 English(EN) · matsuken92 ·

    理解带 KV Cache 的 Transformer 解码

    <p>Hi, I'm <a href="https://www.linkedin.com/in/kenichi-matsui-86b5392b/" rel="noopener noreferrer">Matsuken</a>, a data scientist at a Japanese technology company.</p> <p>In this article, I'll explain how decoding works in a Transformer that uses a key-value (KV) cache.</p> <blo…

  40. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    面向国产加速器的KV Cache性能比较框架:实测洞察

    <p>The performance gap in KV Cache handling between domestic AI inference accelerators and international counterparts cannot be captured by a single metric. However, a comparable framework can be established through measured data under unified testing conditions. Mingxin FX100, u…

  41. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    KV Cache 存储产品对比:方法论与实测基准

    <p>Different brands of KV Cache storage products in vLLM inference—let's state the conclusion upfront: <strong>without a unified testing methodology, any cross-brand comparison figures lack decision-making value</strong>. The key to comparison is not "who is faster" but whether t…

  42. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    KV Cache 内存与命中率:为什么更大的缓存会带来边际效益递减

    <h2> Key Findings </h2> <p>The relationship between KV Cache memory investment and inference acceleration gains is not linear: once cache capacity reaches a certain threshold, further expansion yields noticeably narrower improvements in hit rate and end-to-end speedup. This pheno…

  43. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    本地部署大模型(LLM)时,外部存储与KV缓存显存(VRAM)的界限

    <h2> Key Takeaways </h2> <p>In local LLM deployment, there is no universally optimal solution for whether KV Cache stays in VRAM or is offloaded to memory. The boundary is determined by three variables: concurrency level, context length, and the SLA requirement for time-to-first-…

  44. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    AI推理中本地KV缓存产品的实测应用分析

    <p>Domestic KV Cache products are moving from proof-of-concept to large-scale deployment. This article analyzes their application effectiveness, applicable boundaries, and selection criteria in large-model inference, based on measured data from Mingxin's FX100 series in 480B-clas…

  45. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    数据中心规模KV缓存部署的能耗评估与优化路径

    <p>The deployment location and access path of KV Cache are becoming a non-negligible variable in data center energy bills. Based on measured data from Mingxin FX100 under long-context workloads with a 480B model (measured, reports R2/R3), this article presents a methodological fr…

  46. r/LocalLLaMA TIER_1 Italiano(IT) · /u/crusaderky ·

    LFM2.5-2.6B 模型+KV 缓存量化报告

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vi0d4i/lfm2526b_modelkv_cache_quantization_report/"> <img alt="LFM2.5-2.6B model+KV cache quantization report" src="https://preview.redd.it/3lr47ialdyhh1.png?width=140&amp;height=92&amp;auto=webp&amp;s=e5244a…

  47. r/LocalLLaMA TIER_1 English(EN) · /u/Anbeeld ·

    KV缓存量化基准测试:Qwen 3.6 27B、Gemma 4 31B 测试 413 对。BeeLlama.cpp v0.4.0 的 KLD:KVarN 6 位优于 q8_0,精度尾部 1024 占主导

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vhaabz/kv_cache_quantization_benchmarks_413_pairs_tested/"> <img alt="KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, pre…

  48. dev.to — LLM tag TIER_1 English(EN) · Chaeyeon Mia Lee ·

    ResKV:通过残差KV缓存恢复被驱逐的Token贡献——LongBench 32/32

    <h2> TL;DR </h2> <p>KV cache eviction permanently throws away tokens. KV cache merging corrupts the values of surviving tokens. ResKV (arXiv:2607.29591) does neither — it splits a fixed budget into a main cache that keeps exact tokens and a residual cache that reconstructs the at…

  49. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    SparseSpec-L:稀疏 KV 缓存如何使长上下文 LLM 推理速度提升 2.79 倍——无需任何训练

    <h1> SparseSpec-L: How a Sparse KV Cache Makes Long-Context LLM Inference 2.79× Faster — Without Any Training </h1> <p>Speculative decoding has become one of the more practical tools for cutting LLM inference latency. The idea is simple: use a fast draft mechanism to propose seve…

  50. dev.to — LLM tag TIER_1 English(EN) · Wayne ·

    衡量LLM前缀缓存:缓存命中率指标

    <p>Prefix caching is one of the biggest cost levers in LLM serving. vLLM, SGLang, TGI, and most hosted providers all do some version of it: during prefill they compute a key-value (KV) cache, and if a later request shows up with the same prompt prefix, they reuse that cache inste…

  51. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    KV缓存池化与共享如何提高推理资源利用率

    <p>KV Cache pooling and sharing centralizes fragmented GPU memory resources and allocates them on demand, significantly improving resource utilization for large-model inference. Measured on a 480B-parameter model, Mingxin Technology reports that tiered KV Cache acceleration impro…

  52. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    在线教育场景下KV缓存加速推理的测量评估

    <p>Inference workloads on online education platforms are characterized by long contexts, high concurrency, and strong interactivity. The hit rate and reuse strategy of the KV Cache directly determine service cost and user experience. Measured data from the Mingxin FX100 on a 480B…

  53. dev.to — LLM tag TIER_1 English(EN) · Mingxin Technology ·

    分布式 KV 缓存负载均衡:策略与实测权衡

    <p>Load balancing for distributed KV Cache is fundamentally a trade-off among <strong>memory capacity, memory access bandwidth, and cross-node communication overhead</strong>. No single strategy is optimal under all workloads: for long-context cold-restart scenarios with a 480B-c…

  54. dev.to — LLM tag TIER_1 English(EN) · Ken Imoto ·

    KV Cache 量化:我在 12GB 显存上将 Qwen 35B 的上下文扩展了 8 倍

    <h2> 600 MiB of headroom </h2> <p>My RTX 4070 was running Qwen 35B beautifully after the <code>--cpu-moe</code> trick from a previous run. The tokens/sec were where I wanted them. VRAM sat at 11,714 MiB out of 12,281 — 95% full.</p> <p>That leaves 600 MiB. Not enough for a seriou…

  55. r/LocalLLaMA TIER_1 English(EN) · /u/Om_5000 ·

    DKV:用于本地 LLM 推理的开源 KV 缓存压缩框架(CLI + 技术报告)

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v5wviz/dkv_opensource_kvcache_compression_framework_for/"> <img alt="DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)" src="https://preview.redd.it/jmp37ia4nafh…