PulseAugur
实时 13:11:22
English(EN) DeepSeek V4 Flash on an M2 Ultra: repacked to 141 GiB losslessly, smaller than the Q4 GGUF, at 25.8 t/s (42 t/s peak)

DeepSeek V4 Flash 针对 Apple M2 Ultra 进行优化,速度达 25.8 t/s

一位用户已针对 Apple 的 M2 Ultra 芯片优化了 DeepSeek V4 Flash 模型,实现了显著的性能提升。打包后的模型为 141 GiB,小于公开的 GGUF 版本,运行速度为每秒 25.8 个 token,峰值速度可达每秒 42 个 token。此优化还包括 SSD KV 缓存和动态通道,支持 100 万 token 的上下文窗口。 AI

影响 展示了通过硬件特定优化在本地 LLM 部署中进一步提升性能的潜力。

排序理由 用户驱动的针对特定硬件上现有模型的优化和性能报告。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

DeepSeek V4 Flash 针对 Apple M2 Ultra 进行优化,速度达 25.8 t/s

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Agusx1211 ·

    DeepSeek V4 Flash 在 M2 Ultra 上运行:无损打包至 141 GiB,小于 Q4 GGUF,速度 25.8 t/s(峰值 42 t/s)

    <!-- SC_OFF --><div class="md"><p>This is one more vibe slopped custom optimization for, in this case, my hardware (m2 ultra 60 cores, 192gb). It is just a fork from llama.cpp with a few changes, it achieves:</p> <p>- DeepSeek V4 Flash, no kv cache quant</p> <p>- 141GiB model, by…