A user has detailed their setup for running large language models locally, utilizing six BC-250 boards. Their preferred configuration involves four boards running Qwen Next Flash IQ2_XS with a 100k context window, achieving approximately 28 tokens/sec for short generations and 24 tokens/sec at 50k context. The remaining two boards are used for running a 3.6 35b q4 model with a 100k context window at 60 tokens/sec. This entire setup is managed using llama with Vulkan and RPC over a 1Gb Ethernet connection, with a makeshift cardboard box serving as an intake fan system. AI
IMPACT Demonstrates custom hardware configurations for running LLMs locally, potentially inspiring similar setups for users with specific hardware needs.
RANK_REASON User-built hardware configuration for running LLMs.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →