PulseAugur
EN
LIVE 11:21:45

Flash-MoE enables 397B parameter LLM on 48GB laptop

A new technique called Flash-MoE allows a massive 397 billion parameter model, Qwen3.5-397B-A17B, to run on consumer hardware like a MacBook Pro with only 48GB of RAM. This is achieved by leveraging a Mixture-of-Experts (MoE) architecture, where only a fraction of the model's parameters are activated per token, allowing the rest to be streamed from an SSD. This approach, inspired by a previously unpublished Apple paper, significantly reduces the memory footprint, though it results in slower inference speeds compared to cloud-based solutions. AI

IMPACT Enables running large models on consumer hardware, potentially expanding local AI use cases despite slower speeds.

RANK_REASON Demonstrates a novel technique for running large models on consumer hardware, inspired by a prior research paper.

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Flash-MoE enables 397B parameter LLM on 48GB laptop

How we ranked this

Signal score
27 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Demonstrates a novel technique for running large models on consumer hardware, inspired by a prior research paper.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Dishant Sharma ·

    Flash-MoE: How to Run a 397B Model on a Laptop

    <p>Reddit user Several-Tax31 had one response when Flash-MoE dropped this week. He compared running a 397B model on a laptop to discussing perpetual motion machines. "The second principle of local inference," he wrote, "states that a model needs to fit in RAM and VRAM to run at d…