PulseAugur
实时 15:34:15
English(EN) 48 tg/s 440 prefill on my grandma's cluster (2xP40) (sort of)

双 Tesla P40 设置使用 Qwen 3.8 27B 达到 48 tok/s

一位 Reddit 用户分享了他们优化双 Tesla P40 设置以进行本地 LLM 推理的经验,实现了显著提高的 token 生成速度。通过将 KV 缓存切换到 FP16 并微调 MTP 投机解码等参数,他们观察到短上下文速度高达 48 tokens/s,比之前的 15 tokens/s 有了大幅提升。基准测试结果还显示了令人印象深刻的预填充速度,以 258.33 tokens/s 处理了 63,900 个 token 的前缀。 AI

影响 通过硬件和软件优化,展示了本地 LLM 推理速度显著提升的潜力。

排序理由 用户生成的关于优化 LLM 推理硬件的指南。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

双 Tesla P40 设置使用 Qwen 3.8 27B 达到 48 tok/s

本文如何被排名

Signal score
3 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
用户生成的关于优化 LLM 推理硬件的指南。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Jumpy-Operation-4615 ·

    48 t/s 440 prefill 在我奶奶的集群上(2xP40)(某种程度上)

    <!-- SC_OFF --><div class="md"><p><strong>TL;DR:</strong> switching KV cache to f16 may give a boost in speed if using MTP and ngrams. </p> <p>I have a self-built &quot;AI mega-cluster&quot; with 2x P40s on a cheap Chinese motherboard and a Xeon CPU (around $1,100 to build, inclu…