PulseAugur
EN
LIVE 20:20:24

User trains 3.87B MoE model Apex-2 from scratch, matching Qwen2.5-1.5B on coding

A user has trained a 3.87B parameter Mixture-of-Experts (MoE) model, named Apex-2, from scratch using 86.5 billion tokens for pre-training and an additional 2.5 billion tokens for supervised fine-tuning. The model features a decoder-only MoE architecture with 16 experts and an active parameter count of 1.45 billion per token, supporting a 4096 token context window. While Apex-2 shows competitive performance on coding benchmarks, matching Qwen2.5-1.5B with significantly less pre-training data, it struggles with knowledge-intensive tasks and exhibits frequent hallucinations due to the limited pre-training dataset. AI

IMPACT Demonstrates efficient training of smaller MoE models, potentially enabling more accessible custom model development.

RANK_REASON User-trained model release with detailed technical specifications and benchmark results. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

User trains 3.87B MoE model Apex-2 from scratch, matching Qwen2.5-1.5B on coding

How we ranked this

Signal score
4 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
User-trained model release with detailed technical specifications and benchmark results. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Prestigious-Taste-63 ·

    I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wxiy8y/i_trained_a_387b_moe_145b_active_from_scratch_on/"> <img alt="I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens" src="https://external-preview.redd.it/g-EbVuBrYp4d5WMZy18GUkfblNX58…