PulseAugur
EN
LIVE 13:01:45

Qwen 3.8 27B model performance tested on 12GB and 24GB VRAM systems

Users on Reddit's r/LocalLLaMA are comparing different versions and quantizations of the Qwen large language model to determine optimal performance on consumer hardware. One user details how to achieve a 194,048 token context with Qwen 3.8 27B on a 24GB VRAM system, noting a trade-off between context length and generation speed. Another user tests Qwen3.8 27B at Q2 and Q3 quantizations against Qwen3.6 35B-A3B MoE on a 12GB VRAM laptop, finding the MoE model offers better generation speed and quality for that hardware. AI

IMPACT Provides practical insights for users optimizing LLM performance on limited hardware, influencing model quantization and architecture choices.

RANK_REASON User-driven comparative analysis of open-source LLM performance on consumer hardware.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Qwen 3.8 27B model performance tested on 12GB and 24GB VRAM systems

COVERAGE [2]

  1. r/LocalLLaMA TIER_1 (AF) · /u/sisyphus-cycle ·

    Qwen 3.8 27b in 24gb of VRAM

    <!-- SC_OFF --><div class="md"><p>Thought i would just add my own flags here for llama.cpp (literally pulled and rebuilt latest today). Running on a 4090 FE and 48gb of DDR4 ram on WSL2.</p> <p>Basically you have 3 knobs you can tune for maximum context. I prefer to keep my kv ca…

  2. r/LocalLLaMA TIER_1 Deutsch(DE) · /u/CoffeeToCode99 ·

    Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12GB VRAM

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vq60on/qwen38_27b_q2_vs_q3_vs_qwen36_35ba3b_moe_on_12gb/"> <img alt="Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12GB VRAM" src="https://preview.redd.it/io9im15fgsjh1.png?width=140&amp;height=69&amp;auto=w…