A new technical report introduces A.X K2, a 688 billion parameter Mixture-of-Experts (MoE) language model designed for agentic applications. Despite being trained on fewer tokens than its predecessor, A.X K1, A.X K2 demonstrates significant improvements across various benchmarks due to enhanced token efficiency and a higher-quality training dataset. The model incorporates innovations like Sparse Gated Attention (SGA) for efficient long-context processing and Gated Norm (GN) for stable large-scale training, achieving strong performance on benchmarks like RULER and math tasks. AI
IMPACT Introduces a new large-scale MoE model with innovations for agentic applications and long contexts, potentially advancing LLM efficiency and capabilities.
RANK_REASON Publication of a technical report detailing a new large language model. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- A.X K2
- Fp8
- Gated Norm
- Hugging Face
- mixture of experts
- NVFP4
- RULER
- Sparse Gated Attention
- Think-Fusion
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →