PulseAugur
EN
LIVE 06:06:54

Users benchmark LLMs like Gemma and Qwen on consumer hardware

Users are experimenting with running large language models like Gemma and Qwen on consumer hardware. One user reported achieving 70 tokens per second with Gemma4:12b-it-qat on a Lenovo with an NVIDIA 4070, while Qwen3.8:27b ran at a slower 6 tokens per second. Another user sarcastically noted Qwen3.8's performance of 37 tokens per second on a high-end GPU, questioning the practical value of such experiments. AI

IMPACT Demonstrates the increasing accessibility of running large language models on consumer-grade hardware, enabling broader experimentation.

RANK_REASON User-level benchmarking and experimentation with open-source LLMs on consumer hardware.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Users benchmark LLMs like Gemma and Qwen on consumer hardware

COVERAGE [2]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Today, I got Gemma4:12b-it-qat running at a good 70 tokens per second on the Lenovo, Nvidia 4070, llama.cpp, with multi token prediction. Got Qwen3.8:27b runnin

    Today, I got Gemma4:12b-it-qat running at a good 70 tokens per second on the Lenovo, Nvidia 4070, llama.cpp, with multi token prediction. Got Qwen3.8:27b running at a good 6 tokens per second. Yeah that one ain't going anywhere fast. I'm downloading Gemma4:26b to see how fast it …

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    🚀🎉 Wow, it's truly groundbreaking news that # Qwen3 .8 can churn out a whole 37 tokens per second on a # GPU you'd need a # loan to buy! 💸 Apparently, this "exp

    🚀🎉 Wow, it's truly groundbreaking news that # Qwen3 .8 can churn out a whole 37 tokens per second on a # GPU you'd need a # loan to buy! 💸 Apparently, this "experiment" showed that throwing every kitchen sink at a model doesn't exactly equal genius-level # AI . 🙄 https:// piszcze…