Kivi
PulseAugur coverage of Kivi — every cluster mentioning Kivi across labs, papers, and developer communities, ranked by signal.
-
New technique improves Transformer KV cache compression
Researchers have developed Codec-Gauge, a post-training layer designed to improve the compression of Key-Value (KV) caches in long-context Transformer models. This method learns orthogonal channel transforms that optimi…
-
New techniques aim to improve LLM KV-cache efficiency and accuracy
Researchers are exploring novel methods to improve the efficiency of large language models by optimizing their KV cache, a component crucial for inference but known for its high memory and bandwidth demands. One approac…
-
New KV Cache Compression Techniques Boost LLM Inference Performance · 9 sources tracked
Multiple research papers explore novel techniques for optimizing the Key-Value (KV) cache in large language model (LLM) serving to address memory and performance bottlenecks. These methods, including quantization, pruni…