PulseAugur
实时 04:32:56
English(EN) 🚀🎉 Wow, it's truly groundbreaking news that # Qwen3 .8 can churn out a whole 37 tokens per second on a # GPU you'd need a # loan to buy! 💸 Apparently, this "exp

用户在消费级硬件上对 Gemma 和 Qwen 等大语言模型进行基准测试

用户正在尝试在消费级硬件上运行 GemmaQwen 等大型语言模型。一位用户报告称,在配备 NVIDIA 4070 的联想电脑上,Gemma4:12b-it-qat 的运行速度达到了每秒 70 个 token,而 Qwen3.8:27b 的运行速度较慢,为每秒 6 个 token。另一位用户讽刺地指出 Qwen3.8 在高端 GPU 上的性能为每秒 37 个 token,质疑此类实验的实际价值。 AI

影响 展示了在消费级硬件上运行大型语言模型的可及性日益提高,从而能够进行更广泛的实验。

排序理由 用户在消费级硬件上对开源大语言模型进行基准测试和实验。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

用户在消费级硬件上对 Gemma 和 Qwen 等大语言模型进行基准测试

报道来源 [2]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    今天,我在联想电脑、Nvidia 4070、llama.cpp 上成功运行了 Gemma4:12b-it-qat,速度达到每秒 70 个 token,并支持多 token 预测。同时还运行了 Qwen3.8:27b

    Today, I got Gemma4:12b-it-qat running at a good 70 tokens per second on the Lenovo, Nvidia 4070, llama.cpp, with multi token prediction. Got Qwen3.8:27b running at a good 6 tokens per second. Yeah that one ain't going anywhere fast. I'm downloading Gemma4:26b to see how fast it …

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🚀🎉 哇,#Qwen3.8 竟然能在你得贷款才能买得起的 #GPU 上每秒生成 37 个 token,这消息真是突破性的!💸 据说,这个“exp

    🚀🎉 Wow, it's truly groundbreaking news that # Qwen3 .8 can churn out a whole 37 tokens per second on a # GPU you'd need a # loan to buy! 💸 Apparently, this "experiment" showed that throwing every kitchen sink at a model doesn't exactly equal genius-level # AI . 🙄 https:// piszcze…