A user has successfully configured a cluster of four AMD BC-250 mining boards to run large language models locally. The setup, costing approximately $500, achieves impressive performance metrics, with Next Flash IQ3_XXS reaching up to 70 tokens/sec at a 100k context window and Qwen 3.6 35B A3B IQ4 models running at 145 tokens/sec with a 256k context. The user detailed optimizations made to the system, including speculative decoding, graph replay, and efficient context handling, and noted the system's power consumption and thermal throttling challenges. AI
IMPACT Demonstrates cost-effective hardware solutions for running large language models locally.
RANK_REASON User-driven repurposing of mining hardware for LLM inference.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →