PulseAugur
EN
LIVE 10:26:11
ENTITY KV caches

KV caches

PulseAugur coverage of KV caches — every cluster mentioning KV caches across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
1
4 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
1
3 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/1 · 4 TOTAL
  1. TOOL · CL_190063 ·

    BinaryPC offers training-free sparse attention for efficient LLM decoding

    Researchers have developed BinaryPC, a novel sparse attention mechanism designed to improve the efficiency of long-context large language models. This training-free method uses binary principal components to create comp…

  2. TOOL · CL_126863 ·

    Llama-server bug discarded KV caches, fix restores fast state restores

    A bug in llama-server caused it to discard restored KV caches, forcing a full re-prefill and significantly increasing processing time. The issue stemmed from the server's state saving mechanism, which serialized token d…

  3. TOOL · CL_119710 ·

    InfoFlow KV improves retrieval-augmented generation for long contexts

    Researchers have developed InfoFlow KV, a novel method for improving retrieval-augmented generation (RAG) in large language models. This technique addresses the bottleneck of prefilling large retrieved contexts during i…

  4. RESEARCH · CL_05362 ·

    TurboQuant compresses AI vectors to 2-4 bits without accuracy loss

    A new method called TurboQuant has been developed to compress AI vectors, such as those in KV caches and attention keys, to as few as 2-4 bits per number without sacrificing accuracy. This technique relies on the princi…