Kimi K3 has released its technical report detailing significant architectural innovations aimed at improving the efficiency and scalability of large language models, particularly for long-context tasks and agentic operations. The model introduces Kimi Delta Attention (KDA) to manage long sequences by using a compressed, recursive state instead of full KV cache, and Gated Multi-Layer Attention (MLA) for periodic checks of the full context. Attention Residuals (AttnRes) allow deeper networks to selectively access information from earlier layers, addressing information compression issues in deep models. For width expansion, LatentMoE compresses hidden states before expert processing, reducing communication overhead, while SiTU-GLU mitigates activation overflow risks. The training system also incorporates partial rollouts and token-level regularization for reinforcement learning, enabling agents to handle long, asynchronous tasks by pausing and resuming execution, and preserving external environment states for seamless task recovery. AI
IMPACT Introduces novel architectural components that could improve efficiency and scalability for long-context processing and agentic tasks.
RANK_REASON Frontier-lab model release with system card [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Attention Residuals
- Innu-aimun
- Kimi Delta Attention
- Kimi K2
- Kimi K3
- LatentMoE
- SiTU-GLU
- SwiGLU
- Transformer++
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →