PulseAugur
EN
LIVE 21:35:45

Mimo 2.5 Pro hits 83 t/s on Nvidia GB10 cluster

The Mimo 2.5 Pro large language model has been benchmarked on an 8x Nvidia GB10 cluster, achieving impressive throughput speeds. Under single-user conditions, it reached 40 tokens/second with a 1k context, scaling up to 17 tokens/second with a 250k context. With parallel processing, the model demonstrated even higher performance, hitting 83 tokens/second with four parallel requests. AI

IMPACT Demonstrates high throughput for large context windows on specialized hardware, potentially influencing local LLM deployment strategies.

RANK_REASON Benchmark results for a specific model on custom hardware. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Mimo 2.5 Pro hits 83 t/s on Nvidia GB10 cluster

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Benchmark results for a specific model on custom hardware. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
121 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/ciprianveg ·

    Mimo 2.5 Pro - 40t/s on 8x Nvidia Spark/GB10 cluster

    <!-- SC_OFF --><div class="md"><p>I got Mimo 2.5 Pro running on my 8x Asus Nvidia GB10 cluster using mtp-2, single user request, coding:<br /> 40 t/s - 1k context,<br /> 32t/s - 30k context,<br /> 25t/s - 125k context,<br /> 17t/s - 250k context.</p> <p>2 parallel reached 60t/s a…