A new technique called Flash-MoE allows a massive 397 billion parameter model, Qwen3.5-397B-A17B, to run on consumer hardware like a MacBook Pro with only 48GB of RAM. This is achieved by leveraging a Mixture-of-Experts (MoE) architecture, where only a fraction of the model's parameters are activated per token, allowing the rest to be streamed from an SSD. This approach, inspired by a previously unpublished Apple paper, significantly reduces the memory footprint, though it results in slower inference speeds compared to cloud-based solutions. AI
IMPACT Enables running large models on consumer hardware, potentially expanding local AI use cases despite slower speeds.
RANK_REASON Demonstrates a novel technique for running large models on consumer hardware, inspired by a prior research paper.
- Apple Inc.
- CVS Health
- Dan Woods
- Flash MoE
- GitHub
- Hacker News
- LLM in a Flash: Efficient Large Language Model Inference with Limited Memory
- MacBook Pro M3 Max
- Qwen3.5-397B-A17B
- r/LocalLLaMA
- Several-Tax31
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →