mooncake
PulseAugur coverage of mooncake — every cluster mentioning mooncake across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
AI infrastructure evolves to integrate storage for LLM inference
The AI infrastructure landscape is shifting from solely focusing on GPU compute to a more integrated approach involving compute, networking, memory, and storage. This evolution is driven by the demands of large language…
-
Mingxin FX100 storage solution accelerates video inference, reducing latency
Mingxin's FX100 storage solution addresses latency bottlenecks in real-time video inference, which are often caused by storage and data path limitations rather than GPU compute. The system employs a tiered KV cache appr…
-
New research optimizes disaggregated LLM inference with topology-aware data movement
A new research paper introduces a topology-aware data movement system designed to optimize disaggregated LLM inference. The system addresses the challenge of transferring KV caches between separate GPU pools by discover…
-
Moonshot AI open-sources Kimi K3, valuation hits $31.5B · 2 sources tracked
Moonshot AI has fully open-sourced its Kimi K3 model, a 2.8T parameter model, and revealed the 401 contributors behind it. This release coincides with the company's valuation soaring to $31.5 billion, with each employee…
-
Moonshot AI launches Kimi K3 with 1M context, powering Agentic Slides · 1 source tracked
Moonshot AI has launched Kimi K3, a new flagship model with 2.8 trillion parameters and a 1 million token context window. This model powers the Kimi Agentic Slides tool, which generates presentations from various docume…
-
Qingjing Technology establishes East China HQ, plans 10k-card AI Token factory
Qingjing Technology, a company specializing in AI Token production services, has established its East China regional headquarters in Qianjiang Century City, Hangzhou. The company plans to build a high-quality AI Token f…
-
Alibaba's Qwen3.5 hits record 580 tps for agentic workloads
Alibaba's Qwen team has achieved a new record for agentic workloads, reaching 580 trillion tokens per second on the TokenSpeed engine. This significant performance boost was accomplished with the help of several key par…
-
Moore Threads rallies open-source AI dev community for MUSA GPU ecosystem
Chinese GPU maker Moore Threads has convened a meetup focused on integrating its MUSA architecture with key open-source large model inference frameworks like SGLang. The event brought together core developers from proje…