PulseAugur
EN
LIVE 10:38:45
ENTITY Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

PulseAugur coverage of Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — every cluster mentioning Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
1
1 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
0 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

1 day(s) with sentiment data

RECENT · PAGE 1/1 · 1 TOTAL
  1. TOOL · CL_206769 ·

    GPU idle time and fragmentation slash LLM inference throughput

    GPU idle time and memory fragmentation significantly reduce inference throughput in large language model serving, often masked by compute utilization metrics. Research, including work on vLLM's PagedAttention and findin…