A developer has successfully trained a miniature version of the Kimi K3 model, named "mini Kimi-K3," for under $250. This new model, featuring Kimi K3's architecture including Kimi Delta Attention and LatentMoE, boasts 1.02 billion parameters and was trained on approximately 5 billion tokens. Despite its small scale and lack of instruction tuning, the mini Kimi-K3 has demonstrated superior performance on the HellaSwag benchmark, achieving a 33.4% score that surpasses the GPT-2 124M model's score of 28%. AI
IMPACT Demonstrates cost-effective methods for replicating advanced model architectures, potentially lowering barriers to AI research and development.
RANK_REASON The cluster describes the creation and evaluation of a new, smaller-scale AI model based on an existing architecture, including benchmark results. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →