IFM has released the K2-Horizon-MoVA-36B-A4B model, a sparse Mixture-of-Experts model that utilizes Mixture-of-Values attention. This model boasts 36 billion parameters but only activates 4 billion per token, achieving frontier-class results that outperform larger dense and MoE models on agentic and reasoning benchmarks. It also supports a 512K context window and will have its intermediate checkpoints, training data, and code released to the public. AI
IMPACT This sparse MoE model demonstrates competitive performance with significantly fewer active parameters, potentially influencing future efficient model architectures.
RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →