PulseAugur
EN
LIVE 18:34:45

124B model achieves 38.7 tok/s on single DGX Spark, outperforming DeepSeek V4 Flash

A user named sudoingX benchmarked a 124B parameter model on a single DGX Spark, achieving 38.7 tokens/second on the optimized INT4 path. This performance was found to be 2.4 times faster than DeepSeek V4 Flash on the same hardware. The user initially reported issues with the official quantizations on a single Spark but later corrected this, confirming it as their fastest option. AI

IMPACT Demonstrates significant performance gains for large models on consumer-grade hardware, potentially influencing hardware choices and model optimization strategies.

RANK_REASON User-conducted benchmark of a specific model's performance on hardware, comparing it to another model. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

124B model achieves 38.7 tok/s on single DGX Spark, outperforming DeepSeek V4 Flash

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/AcanthisittaOk1699 ·

    Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vmj6a3/benched_a_124b_on_one_dgx_spark_for_a_week_and/"> <img alt="Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the s…