PulseAugur
实时 19:13:07
English(EN) Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box

124B 模型在单台 DGX Spark 上实现 38.7 tok/s,性能优于 DeepSeek V4 Flash

一位名为 sudoingX 的用户在一台 DGX Spark 上对一个拥有 124B 参数的模型进行了基准测试,在优化的 INT4 路径上实现了每秒 38.7 个 token 的速度。该性能被发现比在同一硬件上运行的 DeepSeek V4 Flash 快 2.4 倍。用户最初报告了在单台 Spark 上使用官方量化版本时遇到的问题,但后来纠正了这一点,确认这是他们最快的选项。 AI

影响 展示了在消费级硬件上运行的大型模型的显著性能提升,可能影响硬件选择和模型优化策略。

排序理由 用户对特定模型在硬件上的性能进行的基准测试,并与其他模型进行了比较。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

124B 模型在单台 DGX Spark 上实现 38.7 tok/s,性能优于 DeepSeek V4 Flash

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/AcanthisittaOk1699 ·

    Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vmj6a3/benched_a_124b_on_one_dgx_spark_for_a_week_and/"> <img alt="Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the s…