PulseAugur
中
实时 22:39:20
English(EN) Strata on a power limited 5090 and 96GB of DDR5-6400 is cranking out 150-200 tok/s decode and 5-6k prefill! Qwen3.8-Flash-Next at IQ3_S, CTX at 128k tokens (8-bit).

Qwen3.8-Flash-Next 在本地硬件上实现 200 tok/s 解码速度

一位 Reddit 用户报告了 Qwen3.8-Flash-Next 模型在功耗受限的配置下取得了令人印象深刻的性能指标。使用 Strata 软件,在配备 96GB DDR5 内存的 5090 GPU 上,系统实现了 150-200 tokens/秒 的解码速度和 5-6k tokens 的预填速度。该配置还支持 128k tokens 的上下文窗口,量化级别为 IQ3_S。 AI

影响 展示了大型语言模型的高效本地部署,可能降低高级人工智能使用的门槛。

排序理由 用户关于特定模型和软件配置在本地硬件上性能的报告。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Qwen3.8-Flash-Next 在本地硬件上实现 200 tok/s 解码速度

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
用户关于特定模型和软件配置在本地硬件上性能的报告。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/z0_o6 ·

    在功耗受限的 5090 和 96GB DDR5-6400 上,解码速度达到 150-200 tok/s,预填充速度达到 5-6k!Qwen3.8-Flash-Next 在 IQ3_S 下,上下文长度为 128k tokens (8-bit)。

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wvwssq/strata_on_a_power_limited_5090_and_96gb_of/"> <img alt="Strata on a power limited 5090 and 96GB of DDR5-6400 is cranking out 150-200 tok/s decode and 5-6k prefill! Qwen3.8-Flash-Next at IQ3_S, CTX at 1…