PulseAugur
EN
LIVE 14:33:40
(AF) Could we have a --disk-moe or --n-disk-moe like --cpu-moe or --n-cpu-moe so we can use disk/cpu/gpu ?

Reddit user proposes disk-based MoE offloading for larger models

A user on the r/LocalLLaMA subreddit is proposing new features for Mixture of Experts (MoE) models, specifically suggesting "--disk-moe" or "--n-disk-moe" options. These would allow for offloading model layers to disk or central processing units (CPUs) in addition to graphics processing units (GPUs), enabling larger models to run on less powerful hardware. AI

IMPACT Enables running larger models on consumer hardware by leveraging disk storage for offloading.

RANK_REASON User proposal on a subreddit about local LLM deployment.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Reddit user proposes disk-based MoE offloading for larger models

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 (AF) · /u/storm1er ·

    Could we have a --disk-moe or --n-disk-moe like --cpu-moe or --n-cpu-moe so we can use disk/cpu/gpu?

    <!-- SC_OFF --><div class="md"><p>Explicit title, It would be nice to have the ability to have 3 tiers moe offload :( </p> </div><!-- SC_ON --> &#32; submitted by &#32; <a href="https://www.reddit.com/user/storm1er"> /u/storm1er </a> <br /> <span><a href="https://www.reddit.com/r…