PulseAugur
EN
LIVE 09:26:08
中文(ZH) 实测小米最快1T大模型:吞吐量每秒1000+ Tokens,Vibe Coding七秒交付

Xiaomi achieves 1000 tokens/sec on 1T-parameter model with commodity GPUs

Xiaomi's MiMo team has released MiMo-V2.5-Pro-UltraSpeed, a new inference mode for their 1-trillion-parameter model that achieves over 1000 tokens per second on commodity GPUs. This significant speedup is attributed to a combination of FP4 quantization, DFlash speculative decoding, and the TileRT serving system, without requiring custom hardware. The company claims this advancement will revolutionize AI applications by enabling faster parallel reasoning, improving coding agent efficiency, and supporting real-time decision-making processes. AI

IMPACT Accelerates real-time AI applications and agentic workflows by drastically reducing inference latency on widely available hardware.

RANK_REASON This is a significant advancement in AI inference speed and efficiency, demonstrating a new capability on commodity hardware.

Read on 量子位 (QbitAI) →

AI-generated summary · Google Gemini · from 10 sources. How we write summaries →

Xiaomi achieves 1000 tokens/sec on 1T-parameter model with commodity GPUs

COVERAGE [10]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 克雷西 ·

    Actual test of Xiaomi's fastest 1T large model: throughput over 1000 Tokens per second, Vibe Coding delivered in seven seconds

    通用GPU就能实现

  2. 36氪 (36Kr) TIER_1 中文(ZH) ·

    Xiaomi Launches MiMo-V2.5-Pro-UltraSpeed Mode

    36氪获悉,6月8日晚,小米MiMo技术团队正式上线Xiaomi MiMo-V2.5-Pro-UltraSpeed模式。据了解,MiMo-V2.5-Pro-UltraSpeed通过对模型推理系统的全链路工程能力优化,在不降低模型能力前提下,首次把推理速度提升至1000 tokens/s,且无需定制芯片、只使用通用GPU即可达成。

  3. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Xiaomi MiMo and TileRT Push a 1-Trillion-Parameter Model Past 1000 Tokens Per Second on Commodity GPUs

    <p>Xiaomi's MiMo team, with TileRT, released MiMo-V2.5-Pro-UltraSpeed, a serving mode for the MiMo-V2.5-Pro model. It decodes over 1000 tokens per second on a 1-trillion-parameter model using a single 8-GPU commodity node.</p> <p>The post <a href="https://www.marktechpost.com/202…

  4. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    Xiaomi released MiMo-V2.5-Pro-UltraSpeed on June 8, pushing its AI model past 1,000 tokens per second on standard hardware. It runs 15x faster than ChatGPT, cut

    Xiaomi released MiMo-V2.5-Pro-UltraSpeed on June 8, pushing its AI model past 1,000 tokens per second on standard hardware. It runs 15x faster than ChatGPT, cutting complex agentic workflow times from 20 minutes to under 5. # TechNews # Xiaomi # AI # MachineLearning # ChatGPT # O…

  5. Mastodon — fosstodon.org TIER_1 Українська(UK) · [email protected] ·

    Xiaomi Sets New AI Record: Over 1,000 Tokens Per Second on 1 Trillion Parameter Model # # AI # FP4 # GenerativeAI # LLM # MiMoV25Pro # TileRT # UltraSpeedAP

    Xiaomi встановлює новий рекорд AI: понад 1 000 токенів за секунду на 1-трильйонній моделі # # AI # FP4 # GenerativeAI # LLM # MiMoV25Pro # TileRT # UltraSpeedAPI # Xiaomi # XiaomiNews https:// gizchina.net/2026/06/10/xiaomi -mimo-v2-5-pro-1000-tokens-ai-rekord/

  6. Mastodon — fosstodon.org TIER_1 Українська(UK) · [email protected] ·

    Xiaomi Sets New AI Record: Over 1,000 Tokens Per Second on 1 Trillion Parameter Model # # AI # FP4 # GenerativeAI # LLM # MiMoV25Pro # TileRT # UltraSpeedAP

    Xiaomi встановлює новий рекорд AI: понад 1 000 токенів за секунду на 1-трильйонній моделі # # AI # FP4 # GenerativeAI # LLM # MiMoV25Pro # TileRT # UltraSpeedAPI # Xiaomi # XiaomiNews https:// gizchina.net/2026/06/10/xiaomi -mimo-v2-5-pro-1000-tokens-ai-rekord/

  7. Mastodon — fosstodon.org TIER_1 日本語(JA) · [email protected] ·

    📝 'Democratization of Inference Speed' Eliminates AI Developer Disparities — MiMo-V2.5-Pro-UltraSpeed Shows New Competitive Axis in the Era of Ultra-Large Models. Xiaomi announces MiMo-V2.5-Pro-UltraSpeed, capable of running trillion-parameter class models at over 1000 tokens/sec. Open-sourcing foundation models creates a new '...

    📝 「推論速度の民主化」がAI開発者の格差を解消する——MiMo-V2.5-Pro-UltraSpeedが示す、超大規模モデル時代の新しい競争軸 Xiaomiが1兆パラメーター級モデルを1000トークン/秒超で動かすMiMo-V2.5-Pro-UltraSpeedを発表。基盤モデルのオープンソース化により、AI開発の「推論コスト格差」という暗黙の構造が崩壊し始めている。 🔗 https:// techscope365.com/1033/ # 大規模言語モデル # 推論最適化 # オプンソスAI # AI # テクノロジー

  8. Mastodon — mastodon.social TIER_1 日本語(JA) · [email protected] ·

    "MiMo-V2.5-Pro-UltraSpeed" appears, running a 1 trillion parameter class model at an explosive speed of over 1000 tokens/second, with the foundation model released as open source https://fed.brid.gy/r/https://gigazine.net/news/20260609-xiaomi-v2-5-pro-ultra

    1兆パラメーター級モデルを1000トークン/秒超という爆速で動かす「MiMo-V2.5-Pro-UltraSpeed」が登場、基盤モデルはオープンソースで公開 https:// fed.brid.gy/r/https://gigazine .net/news/20260609-xiaomi-v2-5-pro-ultraspeed/

  9. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    Xiaomi's MiMo team has achieved over 1000 tokens per second on a 1-trillion-parameter model using commodity GPUs. The breakthrough comes from extreme model-syst

    Xiaomi's MiMo team has achieved over 1000 tokens per second on a 1-trillion-parameter model using commodity GPUs. The breakthrough comes from extreme model-system codesign combining FP4 quantisation, DFlash speculative decoding and TileRT serving on a single 8-GPU node. https://w…

  10. Mastodon — mastodon.social TIER_1 English(EN) · ngate ·

    🚀 Xiaomi's MiMo-v2.5-Pro-UltraSpeed model is here to redefine "fast" with a staggering 1 trillion parameters and a blazing 1000 TPS, because who doesn't need th

    🚀 Xiaomi's MiMo-v2.5-Pro-UltraSpeed model is here to redefine "fast" with a staggering 1 trillion parameters and a blazing 1000 TPS, because who doesn't need their # AI to outpace their Internet connection? 🤖💨 Now you too can experience the thrill of collaborating with a model th…