Xiaomi has released its MiMo-V2.5 model, notable for its "native omnimodal" architecture. Unlike many multimodal models that integrate separate components for different data types, MiMo-V2.5 was designed from the ground up as a unified system capable of processing text, images, video, and audio simultaneously. This Mixture-of-Experts model boasts approximately 310 billion total parameters with 15 billion active per token, and supports a context window of up to 1 million tokens. AI
IMPACT This native omnimodal architecture could set a new standard for processing diverse data types, potentially improving agent capabilities and long-context reasoning.
RANK_REASON New model release from a major tech company (Xiaomi) with novel architectural claims ('native omnimodal'). [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →