PulseAugur
中
实时 22:55:21
English(EN) Energy Efficiency of Locally Deployed LLMs: A Preliminary Quantitative GPU Power Benchmark on Consumer Hardware

消费级GPU上大语言模型的能效因架构而异,而非仅取决于大小

一篇新论文对本地部署的大语言模型(LLMs)在消费级硬件上的能效进行了基准测试,结果表明,除了参数数量之外,模型架构和量化等因素也显著影响能耗。研究发现,qwen2.5:0.5b和tinyllama:1.1b等小型模型能效最高,而像7B-Mistral这样的大型模型每token消耗的功耗则要高得多。该研究还强调了区分不同token生成模式的重要性,因为有些模型由于内部推理时间过长而表现出异常高的能耗。 AI

影响 强调了本地部署大语言模型在模型大小、架构和能耗之间的权衡,为硬件和软件选择提供指导。

排序理由 该聚类包含一篇详细介绍大语言模型能效量化基准测试的研究论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

消费级GPU上大语言模型的能效因架构而异,而非仅取决于大小

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Philipp M. Z\"ahl, Elja Dalipaj, Anika Hennig, Timon Bayer ·

    本地部署大语言模型的能效:消费级硬件上的初步量化GPU功耗基准测试

    arXiv:2608.00008v2 Announce Type: replace Abstract: The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference. However, the energy costs on consumer hardware remain poorly characterized, as most benchm…

  2. dev.to — LLM tag TIER_1 English(EN) · Ivan Hopkins ·

    大语言模型可移植性演练:在第二个 GPU 云上重新部署一个开放模型

    <p>Access to open weights gives a team the right to operate a model in its own environment. I test portability by asking whether the team can reconstruct the complete service after the host, runtime, or provider changes.</p> <p>During an urgent move, a repository may point to a n…