PulseAugur
中
实时 16:48:25

Mimo 2.5 Pro 在 Nvidia GB10 集群上达到 83 t/s

Mimo 2.5 Pro 大型语言模型已在 8x Nvidia GB10 集群上进行了基准测试,达到了令人印象深刻的吞吐速度。在单用户条件下,其 1k 上下文的吞吐量为 40 tokens/秒,在 250k 上下文下可扩展至 17 tokens/秒。通过并行处理,该模型展示了更高的性能,在四次并行请求下达到了 83 tokens/秒。 AI

影响 在专用硬件上展示了大型上下文窗口的高吞吐量,可能影响本地 LLM 部署策略。

排序理由 特定模型在定制硬件上的基准测试结果。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Mimo 2.5 Pro 在 Nvidia GB10 集群上达到 83 t/s

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
特定模型在定制硬件上的基准测试结果。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
129 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/ciprianveg ·

    Mimo 2.5 Pro - 40t/s 在 8x Nvidia Spark/GB10 集群上

    <!-- SC_OFF --><div class="md"><p>I got Mimo 2.5 Pro running on my 8x Asus Nvidia GB10 cluster using mtp-2, single user request, coding:<br /> 40 t/s - 1k context,<br /> 32t/s - 30k context,<br /> 25t/s - 125k context,<br /> 17t/s - 250k context.</p> <p>2 parallel reached 60t/s a…