PulseAugur
EN
LIVE 23:45:47

Qwen3.8-27B model achieves 21.23 tok/s on Strix Halo hardware

A user shared performance metrics for the Qwen3.8-27B model running with Q4 quantization on Strix Halo hardware. The benchmark showed a median inference speed of 21.23 tokens per second across 96 requests, with a maximum sequence length (n-max) of 4. This data was presented in a Multi Token Prediction matrix. AI

IMPACT Provides specific performance data for a particular model and hardware setup, useful for users evaluating similar configurations.

RANK_REASON User-shared performance metrics for a specific model and hardware configuration.

Read on r/cursor →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3.8-27B model achieves 21.23 tok/s on Strix Halo hardware

COVERAGE [1]

  1. r/cursor TIER_2 English(EN) · /u/creativenew ·

    Qwen3.8-27B Q4 on Strix Halo (395): n-max=4 median 21.23 tok/s, 0 replies under 15 — 96-request MTP matrix

    <table> <tr><td> <a href="https://www.reddit.com/r/cursor/comments/1vwkhzd/qwen3827b_q4_on_strix_halo_395_nmax4_median_2123/"> <img alt="Qwen3.8-27B Q4 on Strix Halo (395): n-max=4 median 21.23 tok/s, 0 replies under 15 — 96-request MTP matrix" src="https://preview.redd.it/3bd5v5…