PulseAugur
实时 01:06:24
English(EN) For Local, what are your minimum good or usable tokens per second, for both promp processing and text generation?

LocalLLaMA 用户讨论最小可用 LLM 性能指标

r/LocalLLaMA 子版块的用户正在讨论在本地运行大型语言模型的最低可接受性能指标。参与者正在分享他们对每秒 token 数的提示处理 (PP) 和文本生成 (TG) 的阈值。一位用户报告称,需要至少 300-350 tokens/秒的提示处理和 9 tokens/秒的文本生成才能认为本地设置可用。 AI

影响 为在消费级硬件上运行 LLM 的用户定义了基线性能预期。

排序理由 用户讨论本地 LLM 部署的性能指标。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LocalLLaMA 用户讨论最小可用 LLM 性能指标

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/panchovix ·

    对于本地模型,在提示处理和文本生成方面,每秒的最小可用或可用 token 数是多少?

    <!-- SC_OFF --><div class="md"><p>Hello guys, hoping you're doing fine.</p> <p>Lately with all the new models, and how popular is offloading, what are your min good or usable t/s for both PP and TG?</p> <p>Speaking on my case, I think PP about 300-350t/s for min, and for TG, abou…