PulseAugur
EN
LIVE 04:57:10

Qwen3.8 27B model hits 280 tok/s with new MXFP4 optimization

A developer has achieved significant performance gains with the Qwen3.8 27B model by implementing MXFP4 kernels on dual R9700 GPUs. This optimization, which utilizes W4A8 quantization, has reportedly surpassed FP8 performance and pushed the hardware to its apparent limits. The developer has open-sourced their work, enabling community collaboration and further development in this area. AI

IMPACT Demonstrates advanced optimization techniques for running large language models on consumer hardware, potentially lowering barriers to entry for local LLM deployment.

RANK_REASON Developer shares optimization techniques for a specific LLM on consumer hardware, including performance metrics and open-sourced code. [lever_c_demoted from research: ic=1 ai=0.7]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3.8 27B model hits 280 tok/s with new MXFP4 optimization

How we ranked this

Signal score
2 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Developer shares optimization techniques for a specific LLM on consumer hardware, including performance metrics and open-sourced code. [lever_c_demoted from research: ic=1 ai=0.7]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/whodoneit1 ·

    How I got 280 tok/s on Qwen3.8 27B on 2xr9700's and 940k tokens kv cache

    <!-- SC_OFF --><div class="md"><p>2 Months ago I had made a post how I was working on my dual R9700's. It's wild to look back at where we were then and where things now stand.</p> <p>Since then after many users commenting and complaining about developers doing the same thing. I t…