PulseAugur
实时 05:47:41
English(EN) Ling-3.0-flash quant ladder on one DGX Spark: the whole thing sits in a 32 to 40 tok/s band

Ling-3.0-flash 模型在 DGX Spark 上不同量化下的速度范围狭窄

一位 Reddit 用户分享了 Ling-3.0-flash 模型的基准测试结果,突显了其在 DGX Spark 系统上的性能。结果显示,在不同的量化级别下,速度范围狭窄,为 32 到 40 token/秒,其中 Q5_K_M 量化速度最快且接近无损。与同一硬件上的 DeepSeek V4 Flash 相比,Ling-3.0-flash 的性能显著优于后者,速度大约是后者的 2.4 倍。 AI

影响 展示了大型语言模型在专用硬件上跨量化级别的有效性能扩展。

排序理由 特定模型和硬件配置的用户生成基准测试结果。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Ling-3.0-flash 模型在 DGX Spark 上不同量化下的速度范围狭窄

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/AcanthisittaOk1699 ·

    Ling-3.0-flash 量化在单个 DGX Spark 上实现 32 到 40 token/秒 的速度

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vlmun8/ling30flash_quant_ladder_on_one_dgx_spark_the/"> <img alt="Ling-3.0-flash quant ladder on one DGX Spark: the whole thing sits in a 32 to 40 tok/s band" src="https://preview.redd.it/lko1z9fxyrih1.jpg?wi…