Users are experimenting with running large language models like Gemma and Qwen on consumer hardware. One user reported achieving 70 tokens per second with Gemma4:12b-it-qat on a Lenovo with an NVIDIA 4070, while Qwen3.8:27b ran at a slower 6 tokens per second. Another user sarcastically noted Qwen3.8's performance of 37 tokens per second on a high-end GPU, questioning the practical value of such experiments. AI
IMPACT Demonstrates the increasing accessibility of running large language models on consumer-grade hardware, enabling broader experimentation.
RANK_REASON User-level benchmarking and experimentation with open-source LLMs on consumer hardware.
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →