PulseAugur
EN
LIVE 14:28:10

Qwen3.8-27B-int4-AutoRound model with MTP spec decode shared

A user on Reddit has shared a quantized version of the Qwen3.8-27B model, specifically the int4-AutoRound variant, which requires 18GB of VRAM. The user highlights that this version includes working MTP (Multi-Turn Prompting) spec decoding, suggesting improved conversational capabilities or efficiency. AI

IMPACT Quantized models like this enable broader local deployment of LLMs on consumer hardware.

RANK_REASON User-shared quantized model release, not from a frontier lab.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3.8-27B-int4-AutoRound model with MTP spec decode shared

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/BusinessMud9586 ·

    Qwen3.8-27B-int4-AutoRound (18GB) - with working MTP spec decode

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vpvwqh/qwen3827bint4autoround_18gb_with_working_mtp_spec/"> <img alt="Qwen3.8-27B-int4-AutoRound (18GB) - with working MTP spec decode" src="https://external-preview.redd.it/gr3-Pf0EHJ7HvkEOgHbOMyrNzd6HKK83BE…