PulseAugur
EN
LIVE 18:11:00
中文(ZH) 苏神复盘 Kimi K3:896 个专家背后,藏着哪些关键技术取舍?

Kimi K3 leverages 896 experts and hybrid attention for efficient scaling

Kimi K3, a 2.8 trillion parameter model, employs a novel approach to manage its massive scale by activating only 16 out of 896 routing experts per token. This strategy, detailed by researcher Su Jianlin, aims to control computational costs despite increased parameters and context length. The model integrates LatentMoE for reduced expert computation and communication, alongside mechanisms like RMSNorm and SiTU-GLU for numerical stability during low-precision training. Its attention architecture combines KDA for continuous state maintenance with MLA for global retrieval, and a gating mechanism to filter results, all while minimizing structural changes for deployment. AI

IMPACT Introduces novel techniques for scaling large models efficiently, potentially influencing future LLM architectures.

RANK_REASON Model release details from a researcher discussing technical choices. [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on 雷峰网 (Leiphone) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Kimi K3 leverages 896 experts and hybrid attention for efficient scaling

COVERAGE [1]

  1. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Su Shen reviews Kimi K3: Behind 896 experts, what key technology trade-offs are hidden?

    <section style="text-align: center; margin: 0px 16px; line-height: 1.75em; display: block;"><img class="rich_pages wxw-img" src="https://static.leiphone.com/uploads/new/images/20260806/6a748fc003c9c.jpg?imageMogr2/quality/90" style="width: 100%; display: inline-block; text-align:…