A user has successfully configured the Qwen3.8-27B model with vision capabilities and a 451K token KV-cache on a single RTX 5090 GPU. The setup, utilizing vLLM and a specific NVFP4 quantization, achieves an average speed of 120 tokens per second, even with a power limit of 400W. This configuration is noted to be accurate for coding tasks and supports up to three parallel sessions. AI
IMPACT Demonstrates efficient local deployment of large context models on consumer hardware.
RANK_REASON User-level configuration and benchmark of an existing model on consumer hardware.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →