PulseAugur
实时 19:04:24
English(EN) KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.

新的 KLQ 量化方法优化 LLM 位宽分配

一项名为 KLQ 的新研究项目,引入了一种用于量化大型语言模型的无训练方法。该方法测量嵌入空间的不均匀性,并根据 KL 散度测量的不同方向的重要性来最优地分配位宽。与依赖方差或可学习旋转的先前方法不同,KLQ 直接评估破坏特定方向的经验成本。虽然有效,但该方法计算量很大,需要多次前向传播才能量化模型。 AI

影响 这种无训练的量化方法可以通过减小模型大小和计算需求,从而实现更高效的 LLM 部署。

排序理由 该条目描述了一个新的研究项目和量化 LLM 的方法,包括技术细节和与现有方法的比较。 [lever_c_demoted from research: ic=1 ai=1.0]

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 KLQ 量化方法优化 LLM 位宽分配

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Federal-Setting-3014 ·

    KLQ:无训练测量旋转量化。在 W4A4KV4 位上超越所有无训练旋转量化方法。Llama 3.2 1B KLQ 量化版超越 SpinQuant,且在无 GPTQ/LDLQ 舍入的情况下接近 ReSpinQuant。

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vk2n2k/klq_trainingfree_measured_rotation_quantization/"> <img alt="KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B…