PulseAugur
EN
LIVE 07:28:30

Qwen3.8-27B model performance discussed on Reddit

Users on the r/LocalLLaMA subreddit are discussing the performance of the Qwen3.8-27B model, specifically its token output speed. One user reported achieving approximately 30-32 tokens per second on a system with a 3090 GPU, 64 GB RAM, and an AMD 7950x CPU, using the Qwen3.8-27B-heretic-ara model with Q5_K_M GGUF quantization via llama.ccp. AI

IMPACT Provides insights into the real-world performance of open-source models for users running them locally.

RANK_REASON User discussion on Reddit about model performance.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3.8-27B model performance discussed on Reddit

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/CooLittleFonzies ·

    How many tokens/second output are you getting with Qwen3.8-27B?

    <!-- SC_OFF --><div class="md"><p>Trying to get a feel for where I stand. If you can list your relevant hardware and model used, that would be awesome.</p> <p>Here's mine:</p> <p><strong>Model:</strong> Qwen3.8-27B-heretic-ara, Q5_K_M GGUF</p> <p><strong>T/s</strong>: ~30-32 t/se…