PulseAugur
实时 21:41:14
English(EN) Kimi K3 full model running on 16x GB10 cluster at 20+tps

Kimi K3 LLM 在 16 个 NVIDIA GB10 芯片上运行,达到 20+ TPS

Kimi K3 大型语言模型已成功部署,并在由 16 个 NVIDIA GB10 Grace Blackwell Superchips 组成的集群上运行。该设置实现了平均每秒超过 20 个 token 的吞吐量,预填充操作的峰值吞吐量为每秒 38 个 token 和每秒 750 个 token。部署是通过 vLLM 实现的,预计将发布复制此设置的说明。 AI

影响 展示了 Kimi K3 模型在先进硬件上的性能能力,可能影响未来的部署策略。

排序理由 该条目详细说明了特定 LLM 在特定硬件集群上的性能,包括基准测试结果。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Kimi K3 LLM 在 16 个 NVIDIA GB10 芯片上运行,达到 20+ TPS

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/ciprianveg ·

    Kimi K3 full model running on 16x GB10 cluster at 20+tps

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vfl525/kimi_k3_full_model_running_on_16x_gb10_cluster_at/"> <img alt="Kimi K3 full model running on 16x GB10 cluster at 20+tps" src="https://preview.redd.it/x4w1912fyehh1.jpeg?width=640&amp;crop=smart&amp;aut…