Users on the r/LocalLLaMA subreddit are discussing the performance of the Qwen3.8-27B model, specifically its token output speed. One user reported achieving approximately 30-32 tokens per second on a system with a 3090 GPU, 64 GB RAM, and an AMD 7950x CPU, using the Qwen3.8-27B-heretic-ara model with Q5_K_M GGUF quantization via llama.ccp. AI
IMPACT Provides insights into the real-world performance of open-source models for users running them locally.
RANK_REASON User discussion on Reddit about model performance.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →