Kimi K3, a 2.8 trillion parameter model, employs a novel approach to manage its massive scale by activating only 16 out of 896 routing experts per token. This strategy, detailed by researcher Su Jianlin, aims to control computational costs despite increased parameters and context length. The model integrates LatentMoE for reduced expert computation and communication, alongside mechanisms like RMSNorm and SiTU-GLU for numerical stability during low-precision training. Its attention architecture combines KDA for continuous state maintenance with MLA for global retrieval, and a gating mechanism to filter results, all while minimizing structural changes for deployment. AI
IMPACT Introduces novel techniques for scaling large models efficiently, potentially influencing future LLM architectures.
RANK_REASON Model release details from a researcher discussing technical choices. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →