PulseAugur
EN
LIVE 22:55:13

Qwen 3.8-27B model runs with 100K context on 16GB GPU

A user on Reddit's r/LocalLLaMA community shared a detailed guide on how to run the Qwen 3.8-27B model with a 100,000 token context window on a 16GB RX 7800 XT GPU. The setup involves compiling llama.cpp with Vulkan support and specific command-line arguments for optimal performance, including using Q4 quantization and a Q8/Q5 KV cache. This guide aims to demonstrate the feasibility of running large context models on consumer-grade hardware. AI

IMPACT Enables running large context models on consumer hardware, potentially lowering barriers for local AI development.

RANK_REASON User-shared guide on running a specific LLM with a large context window on consumer hardware.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen 3.8-27B model runs with 100K context on 16GB GPU

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Haunting-Stretch8069 ·

    Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide

    <!-- SC_OFF --><div class="md"><p>I'm running Qwen 3.8 27B Q4 XS with ~30 t/s decode (no MTP) at 100k context with q8/q5 KV cache on a 16GB AMD GPU. I wanted to share my setup since many believe Qwen 3.8 27B to be infeasible on 16GB VRAM.</p> <p>Build llama.cpp with Vulkan:</p> <…