A user on Reddit is seeking advice regarding the performance of the Qwen3.8-27B model running locally on a Windows 10 system with dual RTX 3090 GPUs. They are experiencing token generation speeds of 50-65 tokens/sec and are questioning if this is normal for their setup, which includes 80GB of RAM and a specific llama.cpp build. The user has provided detailed information about their hardware, software configuration, and the exact command used to run the model, hoping for guidance on potential optimizations or troubleshooting steps. AI
IMPACT Provides insight into local LLM deployment challenges and performance expectations for users with high-end consumer hardware.
RANK_REASON User-generated content seeking help with specific hardware and software configuration for running an LLM locally.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →