PulseAugur
中
实时 08:11:44

消费级GPU在122B模型上实现高LLM速度,运行速度达37 t/s

Reddit的r/LocalLLaMA子版块上一位用户分享了在消费级硬件上运行大型语言模型的令人印象深刻的基准测试。他们使用Nvidia RTX 4090和RTX 5060 Ti,并将部分层卸载到64GB系统内存,在350亿参数模型(35b a3b)上实现了每秒206个token。更值得注意的是,在类似条件下,一个1220亿参数模型(barium-122)的运行速度达到了每秒37个token,超出了用户的乐观预期。 AI

影响 展示了在消费级硬件上运行大型语言模型的重大进展,可能降低AI实验的门槛。

排序理由 用户生成的消费级硬件运行LLM的基准测试。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

消费级GPU在122B模型上实现高LLM速度,运行速度达37 t/s

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
用户生成的消费级硬件运行LLM的基准测试。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
63 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Dry_Long3157 ·

    4090 + 5060 Ti + 64GB RAM:35B-A3B 上 206 t/s,122B 为 37 t/s

    <!-- SC_OFF --><div class="md"><p>I've been benchmarking a two-card box for a few weeks and I still can't quite get over some of these numbers, so I'm dumping them here.</p> <p><strong>Box:</strong> RTX 4090 (24GB) + RTX 5060 Ti (16GB), i9-13900K, 64GB DDR5. WSL2 with 47GB alloca…