A user on Reddit's r/LocalLLaMA subreddit shared their experience running the Qwen 3.8-27B model with the iq3 xxs quantization on an RTX 3060 GPU. They reported achieving between 10-20 tokens per second, with prompt processing taking around 4-7 minutes. The user also provided detailed command-line arguments and a link to a GitHub repository for others looking to replicate the setup on similar hardware, noting that they had over 1GB of VRAM remaining after loading the model and cache. AI
IMPACT Demonstrates feasibility of running large language models on mid-range consumer GPUs.
RANK_REASON User-generated guide for running a specific model on consumer hardware.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →