PulseAugur
中
实时 19:17:04
English(EN) A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode

Strata 分支将 IBM AC922 推理速度提升至 7,350 tokens/sec

一位开发者对 Strata 推理引擎进行了分支,以优化在配备 POWER9 CPU 和 NVIDIA Tesla V100 GPU 的 IBM AC922 服务器上的性能。这个修改后的 Strata 分支实现了显著提升的推理速度,在提示读取方面达到了每秒 7,350 个 token,在生成方面约为每秒 113 个 token。优化侧重于利用硬件能力,包括 NVLink 带宽和 Tensor Core 利用率,以超越同一系统上的先前基准。 AI

影响 在特定硬件配置上优化了推理性能。

排序理由 这是针对特定硬件的现有推理引擎的分支,而不是新的模型发布或重要的行业事件。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Strata 分支将 IBM AC922 推理速度提升至 7,350 tokens/sec

本文如何被排名

Signal score
5 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是针对特定硬件的现有推理引擎的分支,而不是新的模型发布或重要的行业事件。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/okoyl3 ·

    用于 IBM AC922 的 Strata 分支运行 Qwen3.8-FN UD-Q4_K_XL 可实现高达 7,357 tk/s 的预填充和 113 tk/s 的解码

    <!-- SC_OFF --><div class="md"><p>I forked Strata and worked with Claude code with some heavy changes to it to make it work on an IBM AC922 I have access to. The IBM AC922 is a 2018 era beast with two POWER9 20 core SMT4 CPUs that are connected by NVLink to 4 or 6 NVIDIA Tesla V1…