A user has trained a 3.87B parameter Mixture-of-Experts (MoE) model, named Apex-2, from scratch using 86.5 billion tokens for pre-training and an additional 2.5 billion tokens for supervised fine-tuning. The model features a decoder-only MoE architecture with 16 experts and an active parameter count of 1.45 billion per token, supporting a 4096 token context window. While Apex-2 shows competitive performance on coding benchmarks, matching Qwen2.5-1.5B with significantly less pre-training data, it struggles with knowledge-intensive tasks and exhibits frequent hallucinations due to the limited pre-training dataset. AI
IMPACT Demonstrates efficient training of smaller MoE models, potentially enabling more accessible custom model development.
RANK_REASON User-trained model release with detailed technical specifications and benchmark results. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →