Moonshot AI has released Kimi K3, a 2.8 trillion parameter open-weight Mixture of Experts (MoE) model. This model, featuring Kimi Delta Attention and other architectural innovations, offers improved scaling efficiency and excels at complex tasks like long-horizon coding and agentic workflows. The weights are available on Hugging Face in MXFP4 format, and deployment on AWS is detailed using Amazon SageMaker HyperPod and Amazon EKS, requiring substantial GPU resources like NVIDIA B300 Blackwell Ultra GPUs. AI
IMPACT Sets a new benchmark for open-weight MoE models, potentially accelerating self-hosting and research in large-scale AI.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
Read on AWS Machine Learning Blog →
- Amazon Elastic Kubernetes Service
- Amazon SageMaker HyperPod
- AWS
- Gated Multi Head Latent Attention
- Hugging Face
- Kimi Delta Attention
- Kimi K3
- Mixture of Experts
- Moonshot AI
- MXFP4
- NVIDIA B300 Blackwell Ultra GPUs
- Stable LatentMoE
- vLLM
- Amazon EKS
- Kimi K2
- Mixture of Experts (MoE)
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →