Xiaomi's MiMo team has announced MiMo-V2.5-Pro UltraSpeed, a 1 trillion parameter Mixture-of-Experts model capable of exceeding 1,000 tokens per second. This performance was achieved on a standard 8-GPU server, utilizing techniques like FP4 quantization with QAT, DFlash speculative decoding, and TileRT latency-optimized kernels. The company has made this high-speed model available via their API at a premium price for select users. AI
IMPACT Demonstrates a significant leap in inference speed for large models on standard hardware, potentially lowering the cost and increasing accessibility of high-performance AI.
RANK_REASON Significant performance claim for a large model on commodity hardware, indicating a notable advancement in AI inference speed.
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →