A user on Reddit's r/LocalLLaMA subreddit shared their setup involving two NVIDIA RTX A3000M and A2000M GPUs running within a Network Attached Storage (NAS) device. This configuration achieved a speed of 20 tokens per second with the Qwen3.8-27B-GSQ-RCO-IQ3_S model, utilizing a 163k context window. AI
IMPACT Demonstrates a user-level setup for running large language models locally, showcasing performance metrics for specific hardware and model combinations.
RANK_REASON User-generated content detailing hardware configuration and performance for running a large language model.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →