Edge0-AI has released Edge0, an open-source streaming inference engine designed to run large Mixture-of-Experts (MoE) models on consumer hardware. The engine achieves this by memory-mapping the entire model checkpoint on SSD and only loading the actively used experts into RAM for each token. This approach allows a 35 billion parameter model to operate with a peak memory usage of around 3 GB, though it relies heavily on the OS file cache and can impact SSD longevity with frequent reads. AI
IMPACT Enables running larger MoE models on consumer hardware by optimizing memory usage, potentially lowering the barrier to entry for local LLM deployment.
RANK_REASON This is a release of an inference engine and preview models, not a frontier model release from a major lab.
- Apple Silicon
- Edge0
- Edge0-35B-A3B
- Edge0-8B-A1B
- Edge0-AI
- inclusionAI
- Mac Mini M4 Pro
- Qwen3.5-MoE 35B-A3B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →