PulseAugur
EN
LIVE 09:58:59
ENTITY mooncake

mooncake

PulseAugur coverage of mooncake — every cluster mentioning mooncake across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
12
12 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
2
2 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 12 TOTAL
  1. TOOL · CL_277381 ·

    Prime Intellect launches Prime Inference serving platform for open models

    Prime Intellect has launched Prime Inference, a new serving platform designed for frontier open-source AI models. The platform offers both serverless endpoints for variable demand and reserved capacity for sustained wor…

  2. TOOL · CL_256842 ·

    KV Cache Placement Strategies Explored for LLM Memory Efficiency

    A new research paper explores optimal placement strategies for KV caches across different memory tiers (GPU HBM, CPU DRAM, SSD) to manage scarce GPU memory. The study, conducted using a discrete event simulator, found t…

  3. TOOL · CL_209852 ·

    TPU partners with Mooncake for inference optimization

    SemiAnalysis reports that Tensor Processing Unit (TPU) is collaborating with the open-source inference optimization library Mooncake. This partnership aims to integrate TPU capabilities with Mooncake Store, enhancing pe…

  4. TOOL · CL_206773 ·

    LLM TTFT Reduction Strategies Explored Across Compute, Memory, and Storage

    Reducing the first-token latency (TTFT) of large language models is crucial for user experience and performance. This involves optimizing four key areas: compute, GPU memory, storage, and overall architecture. Technique…

  5. RESEARCH · CL_188615 ·

    AI infrastructure evolves to integrate storage for LLM inference

    The AI infrastructure landscape is shifting from solely focusing on GPU compute to a more integrated approach involving compute, networking, memory, and storage. This evolution is driven by the demands of large language…

  6. TOOL · CL_183926 ·

    Mingxin FX100 storage solution accelerates video inference, reducing latency

    Mingxin's FX100 storage solution addresses latency bottlenecks in real-time video inference, which are often caused by storage and data path limitations rather than GPU compute. The system employs a tiered KV cache appr…

  7. TOOL · CL_178265 ·

    New research optimizes disaggregated LLM inference with topology-aware data movement

    A new research paper introduces a topology-aware data movement system designed to optimize disaggregated LLM inference. The system addresses the challenge of transferring KV caches between separate GPU pools by discover…

  8. SIGNIFICANT · CL_173979 ·

    Moonshot AI open-sources Kimi K3, valuation hits $31.5B · 2 sources tracked

    Moonshot AI has fully open-sourced its Kimi K3 model, a 2.8T parameter model, and revealed the 401 contributors behind it. This release coincides with the company's valuation soaring to $31.5 billion, with each employee…

  9. SIGNIFICANT · CL_164703 ·

    Moonshot AI launches Kimi K3 with 1M context, powering Agentic Slides · 1 source tracked

    Moonshot AI has launched Kimi K3, a new flagship model with 2.8 trillion parameters and a 1 million token context window. This model powers the Kimi Agentic Slides tool, which generates presentations from various docume…

  10. TOOL · CL_158968 ·

    Qingjing Technology establishes East China HQ, plans 10k-card AI Token factory

    Qingjing Technology, a company specializing in AI Token production services, has established its East China regional headquarters in Qianjiang Century City, Hangzhou. The company plans to build a high-quality AI Token f…

  11. SIGNIFICANT · CL_55057 ·

    Alibaba's Qwen3.5 hits record 580 tps for agentic workloads

    Alibaba's Qwen team has achieved a new record for agentic workloads, reaching 580 trillion tokens per second on the TokenSpeed engine. This significant performance boost was accomplished with the help of several key par…

  12. RESEARCH · CL_31391 ·

    Moore Threads rallies open-source AI dev community for MUSA GPU ecosystem

    Chinese GPU maker Moore Threads has convened a meetup focused on integrating its MUSA architecture with key open-source large model inference frameworks like SGLang. The event brought together core developers from proje…