A user on Reddit's r/LocalLLaMA subreddit shared a method for running the Qwen3.8 27B model with a 144K context window on an RTX 3090 GPU using vLLM. The user detailed a process involving Ahead-of-Time (AOT) compilation and an FP8 KV cache to achieve this, noting that Just-in-Time (JIT) compilation might cause Out-of-Memory (OOM) errors. Benchmarks provided show impressive token generation speeds, particularly for prompt processing, with the setup achieving high scores on various tasks like tool calling and instruction following. AI
IMPACT Enables larger context windows for local LLM deployments on consumer GPUs.
RANK_REASON User-shared optimization for running a specific LLM on consumer hardware.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →