PulseAugur
EN
LIVE 15:10:52

User seeks advice on Qwen3.8-27B local LLM performance with dual RTX 3090s

A user on Reddit is seeking advice regarding the performance of the Qwen3.8-27B model running locally on a Windows 10 system with dual RTX 3090 GPUs. They are experiencing token generation speeds of 50-65 tokens/sec and are questioning if this is normal for their setup, which includes 80GB of RAM and a specific llama.cpp build. The user has provided detailed information about their hardware, software configuration, and the exact command used to run the model, hoping for guidance on potential optimizations or troubleshooting steps. AI

IMPACT Provides insight into local LLM deployment challenges and performance expectations for users with high-end consumer hardware.

RANK_REASON User-generated content seeking help with specific hardware and software configuration for running an LLM locally.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

User seeks advice on Qwen3.8-27B local LLM performance with dual RTX 3090s

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 Nederlands(NL) · /u/sugarfreecaffeine ·

    Dual RTX 3090 Qwen3.8-27B Help

    <!-- SC_OFF --><div class="md"><p>I'm new to local LLMs and wondering if my performance looks normal or if I'm doing something wrong. My use case is local agentic coding.</p> <h2>Build</h2> <ul> <li><strong>OS:</strong> Windows 10 - NO WSL</li> <li><strong>CPU:</strong> AMD Ryzen…