A user has successfully adapted their streaming stack to run the Qwen3.8-Next model, achieving faster performance than a dense 27b model on their M5 hardware. The Qwen3.8-Next model, when run in a 3-bit quantized version, demonstrated 150 tps for prefill and 3.6 tps for decoding. This performance surpasses the 70 tps prefill and 3 tps decode of the 4-bit dense 27b model on the same M5 hardware. AI
IMPACT Demonstrates improved performance for open-source models on consumer hardware, potentially lowering barriers to entry for local LLM deployment.
RANK_REASON User benchmark of an open-source model on consumer hardware.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →