PulseAugur
EN
LIVE 23:47:02

AMD BC-250 mining boards repurposed for local LLM inference

A user has successfully configured a cluster of four AMD BC-250 mining boards to run large language models locally. The setup, costing approximately $500, achieves impressive performance metrics, with Next Flash IQ3_XXS reaching up to 70 tokens/sec at a 100k context window and Qwen 3.6 35B A3B IQ4 models running at 145 tokens/sec with a 256k context. The user detailed optimizations made to the system, including speculative decoding, graph replay, and efficient context handling, and noted the system's power consumption and thermal throttling challenges. AI

IMPACT Demonstrates cost-effective hardware solutions for running large language models locally.

RANK_REASON User-driven repurposing of mining hardware for LLM inference.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

AMD BC-250 mining boards repurposed for local LLM inference

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
User-driven repurposing of mining hardware for LLM inference.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Ok-Breadfruit-3523 ·

    Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1x2yv1b/running_next_flash_iq3_xxs_at_70_toks_with_100k/"> <img alt="Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex m…