The Kimi K3 large language model has been successfully deployed and is running on a cluster of 16 NVIDIA GB10 Grace Blackwell Superchips. This setup achieved an average throughput of over 20 tokens per second, with a peak of 38 tokens per second and 750 tokens per second for prefill operations. The deployment was facilitated using vLLM, and instructions for replicating this setup are expected to be published. AI
IMPACT Demonstrates the performance capabilities of the Kimi K3 model on advanced hardware, potentially influencing future deployment strategies.
RANK_REASON The item details the performance of a specific LLM on a particular hardware cluster, including benchmark results. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →