A user shared performance metrics for the Qwen3.8-27B model running with Q4 quantization on Strix Halo hardware. The benchmark showed a median inference speed of 21.23 tokens per second across 96 requests, with a maximum sequence length (n-max) of 4. This data was presented in a Multi Token Prediction matrix. AI
IMPACT Provides specific performance data for a particular model and hardware setup, useful for users evaluating similar configurations.
RANK_REASON User-shared performance metrics for a specific model and hardware configuration.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →