A team of five developers is seeking advice on the optimal hardware configuration for self-hosting large open MoE models like Kimi K2.6 and DeepSeek V4 for agentic coding tasks. The primary dilemma is choosing between a dual GH200 NVL2 system with unified memory or an 8x RTX 6000 Blackwell build offering faster VRAM. The user has tested a single GH200, achieving moderate decode speeds but is concerned about prefill performance and the model partially residing in slower unified memory. They are seeking real-world performance data, particularly decode and prefill numbers under concurrency, to make an informed decision within a $100k-$150k budget. AI
IMPACT Guidance for developers on selecting hardware for local LLM deployment, impacting infrastructure choices for AI teams.
RANK_REASON Discussion about hardware choices for running specific LLMs locally, not a new model release or research.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →