PulseAugur
EN
LIVE 21:24:59

Luce Spark enables 35B MoE models on 16GB GPUs

Luce Spark is a new open-source system that enables large 35 billion parameter Mixture-of-Experts (MoE) models to run on a single 16 GB GPU. It achieves this by intelligently keeping only the currently active experts on the GPU, while the rest reside in system RAM and are swapped in as needed. This approach avoids the performance penalty typically associated with offloading, allowing models that would otherwise not fit to run efficiently. AI

IMPACT Enables running large MoE models on consumer hardware, democratizing access to advanced AI capabilities.

RANK_REASON The cluster describes a new open-source system and methodology for running large MoE models on limited hardware, which is a significant research contribution in the field of efficient AI deployment.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Luce Spark enables 35B MoE models on 16GB GPUs

COVERAGE [2]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/sandropuppo ·

    Luce Spark: a 35B MoE on a 16 GB GPU, without the offload tax

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1u0b3cu/luce_spark_a_35b_moe_on_a_16_gb_gpu_without_the/"> <img alt="Luce Spark: a 35B MoE on a 16 GB GPU, without the offload tax" src="https://preview.redd.it/tg6kpi4vs26h1.png?width=640&amp;crop=smart&amp;a…

  2. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    A 35B MoE on a 16 GB GPU, without the offload tax https://www.lucebox.com/blog/spark # HackerNews # Tech # AI

    A 35B MoE on a 16 GB GPU, without the offload tax https://www.lucebox.com/blog/spark # HackerNews # Tech # AI