PulseAugur
实时 05:04:38
English(EN) How I got 280 tok/s on Qwen3.8 27B on 2xr9700's and 940k tokens kv cache

Qwen3.8 27B模型通过新的MXFP4优化达到280 tok/s

一位开发者通过在双R9700 GPU上实现MXFP4内核,显著提升了Qwen3.8 27B模型的性能。据报道,这项利用W4A8量化的优化已经超越了FP8的性能,并将硬件推向了极限。该开发者已开源其工作,促进了社区协作和该领域的进一步发展。 AI

影响 展示了在消费级硬件上运行大型语言模型的先进优化技术,可能降低了本地部署LLM的门槛。

排序理由 开发者分享了在消费级硬件上针对特定LLM的优化技术,包括性能指标和开源代码。[lever_c_demoted from research: ic=1 ai=0.7]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Qwen3.8 27B模型通过新的MXFP4优化达到280 tok/s

本文如何被排名

Signal score
2 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
开发者分享了在消费级硬件上针对特定LLM的优化技术,包括性能指标和开源代码。[lever_c_demoted from research: ic=1 ai=0.7]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/whodoneit1 ·

    如何在2块r9700上以280 tok/s的速度运行Qwen3.8 27B,并拥有940k token的kv cache

    <!-- SC_OFF --><div class="md"><p>2 Months ago I had made a post how I was working on my dual R9700's. It's wild to look back at where we were then and where things now stand.</p> <p>Since then after many users commenting and complaining about developers doing the same thing. I t…