PulseAugur
EN
LIVE 09:26:09

Xiaomi claims 1000+ tokens/sec on 1T parameter model with 8 GPUs

Xiaomi's MiMo team has announced MiMo-V2.5-Pro UltraSpeed, a 1 trillion parameter Mixture-of-Experts model capable of exceeding 1,000 tokens per second. This performance was achieved on a standard 8-GPU server, utilizing techniques like FP4 quantization with QAT, DFlash speculative decoding, and TileRT latency-optimized kernels. The company has made this high-speed model available via their API at a premium price for select users. AI

IMPACT Demonstrates a significant leap in inference speed for large models on standard hardware, potentially lowering the cost and increasing accessibility of high-performance AI.

RANK_REASON Significant performance claim for a large model on commodity hardware, indicating a notable advancement in AI inference speed.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

Xiaomi claims 1000+ tokens/sec on 1T parameter model with 8 GPUs

COVERAGE [3]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/No-Selection2972 ·

    Xiaomi just claimed 1,000+ tps on a 1T model using a standard 8-GPU server

    <!-- SC_OFF --><div class="md"><p>Just saw Xiaomi MiMo announce <strong>MiMo-V2.5-Pro UltraSpeed</strong>, claiming they broke the <strong>1,000 tokens/sec output barrier on a 1 trillion parameter MoE model</strong>. According to them, they’re doing it on a <strong>single standar…

  2. r/singularity TIER_2 English(EN) · /u/Worldly_Evidence9113 ·

    Xiaomi & TileRT just hit 1,000+ TPS on a 1-Trillion Parameter model… on standard commodity GPUs. It’s over for custom silicon?

    &#32; submitted by &#32; <a href="https://www.reddit.com/user/Worldly_Evidence9113"> /u/Worldly_Evidence9113 </a> <br /> <span><a href="https://v.redd.it/yhhcffbui66h1">[link]</a></span> &#32; <span><a href="https://www.reddit.com/r/singularity/comments/1u270b8/xiaomi_tilert_just…

  3. r/singularity TIER_2 English(EN) · /u/elemental-mind ·

    Xiaomi achieves 1000+t/s on 8x commodity GPU cluster with 1T weights model

    <table> <tr><td> <a href="https://www.reddit.com/r/singularity/comments/1u0o9xn/xiaomi_achieves_1000ts_on_8x_commodity_gpu/"> <img alt="Xiaomi achieves 1000+t/s on 8x commodity GPU cluster with 1T weights model" src="https://external-preview.redd.it/djdiNW5tbHY1NTZoMcwdOZIKfH_YdK…