A user on Reddit's r/LocalLLaMA subreddit detailed the construction of a custom home inference server for approximately $3,000. The build features 128GB of VRAM and 256GB of RAM, utilizing components like four V620 GPUs, an EPYC 7452 processor, and a Huananzhi D12D motherboard. The user reported power consumption ranging from 500-900W and expressed satisfaction with its performance, particularly with the Qwen3.8-next-flash Autoround W4A16 model running on a 128k+ context. AI
IMPACT Enables individuals to run larger models locally, potentially reducing reliance on cloud services.
RANK_REASON User-built hardware for AI inference, not a product release from a major lab.
- 128GB VRAM
- 256GB RAM
- DDR4 SDRAM
- EPYC 7452
- Huananzhi D12D
- Lenovo p620 workstation
- Qwen3.8-next-flash Autoround W4A16
- server
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →