A user on Reddit's r/LocalLLaMA subreddit is seeking advice for building a new PC optimized for running large language models like Qwen 3.8 Next and GLM 3 Flash, as well as ComfyUI for image generation. The user plans to utilize llama.cpp and potentially vLLM, with a focus on sharding models across multiple GPUs and offloading to RAM. They are considering a setup with three internal GPUs and two external GPUs via Thunderbolt 5, aiming for a total of 80GB of VRAM using RTX 5060 Ti 16GB cards and 128GB of system RAM. The user is also inquiring about the feasibility of running Linux on a second PC for this setup and whether their proposed hardware configuration has any obvious flaws. AI
RANK_REASON This is a user seeking personal technical advice on a hardware build for running LLMs, not a news event.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →