PulseAugur
中
实时 22:56:48
English(EN) Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards

AMD BC-250 挖矿板被重新用于本地 LLM 推理

一位用户已成功配置了一组四块 AMD BC-250 挖矿板,用于在本地运行大型语言模型。该设置成本约为 500 美元,实现了令人印象深刻的性能指标,Next Flash IQ3_XXS 在 100k 上下文窗口下可达 70 token/秒,Qwen 3.6 35B A3B IQ4 模型在 256k 上下文下可运行至 145 token/秒。用户详细介绍了对系统进行的优化,包括推测解码、图重放和高效上下文处理,并指出了系统的功耗和热节流挑战。 AI

影响 展示了在本地运行大型语言模型的经济高效的硬件解决方案。

排序理由 用户驱动的挖矿硬件重新用于 LLM 推理。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

AMD BC-250 挖矿板被重新用于本地 LLM 推理

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
用户驱动的挖矿硬件重新用于 LLM 推理。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Ok-Breadfruit-3523 ·

    在约 70 tok/s 的速度下运行 Next Flash IQ3_XXS,上下文长度为 10 万,或在约 145 tok/s 的速度下运行 Qwen 3.6 35B A3B IQ4 的 2 个实例,上下文长度为 25.6 万,所有这些都在价值 500 美元的二手矿卡 BC-250 主板上

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1x2yv1b/running_next_flash_iq3_xxs_at_70_toks_with_100k/"> <img alt="Running Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex m…