PulseAugur
实时 11:47:20
English(EN) LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization

新方法通过先进的 2 位和自适应量化提升 LLM 效率

研究人员开发了新的技术,通过先进的量化方法来提高大型语言模型 (LLM) 的效率。一种名为 SPEAR 的方法侧重于量化后的自适应恢复,以最小的开销减小了低比特和全精度模型之间的质量差距。另一种方法 LC-QAT 引入了一个数据高效的 2 位量化感知训练框架,该框架使用线性约束向量量化,能够用显著更少的数据进行有效训练。这些进展旨在使 LLM 的部署更具成本效益和可及性。 AI

影响 能够更高效、更具成本效益地部署 LLM,有可能提高在消费级硬件上的可及性和性能。

排序理由 arXiv 上发表了两篇详细介绍 LLM 量化新方法的学术论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

新方法通过先进的 2 位和自适应量化提升 LLM 效率

报道来源 [6]

  1. arXiv cs.CL TIER_1 English(EN) · Liza Babaoglu, Shuangyi Chen, Ashish Khisti ·

    使用加法码本对大语言模型进行多比特量化

    arXiv:2606.12876v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed across heterogeneous hardware with varying resource constraints, the ability to adaptively manage the trade-off between performance and efficiency without retraining is cri…

  2. arXiv cs.CL TIER_1 English(EN) · Ashish Khisti ·

    使用加法码本对大型语言模型进行多比特量化

    As large language models (LLMs) are increasingly deployed across heterogeneous hardware with varying resource constraints, the ability to adaptively manage the trade-off between performance and efficiency without retraining is critical. We propose Drop-by-Drop, a novel multi-bitw…

  3. arXiv cs.AI TIER_1 English(EN) · Hongyuan Liu, Yawei Li, Zhiqiang Que, Qinli Yang, Junming Shao, Guosheng Hu ·

    SPEAR:一种后量化误差自适应恢复系统,可实现高效低比特大模型服务

    arXiv:2606.11244v1 Announce Type: cross Abstract: Efficient large language model (LLM) serving is increasingly constrained by deployment cost. Quantization is a key technique for reducing serving cost, yet even state-of-the-art 4-bit quantizers exhibit a noticeable quality gap fr…

  4. arXiv cs.AI TIER_1 English(EN) · Haoyu Wang, Xingyu Yu, Haiyan Zhao, Fengxiang Wang, Xu Han ·

    LC-QAT:通过线性约束向量量化实现 LLM 的数据高效 2 位 QAT

    arXiv:2606.10531v1 Announce Type: cross Abstract: Quantization-aware training (QAT) is essential for extremely low-bit large language models (LLMs). Current QAT methods are mainly based on scalar quantization (SQ), which enables efficient optimization but suffers from severe perf…

  5. arXiv cs.AI TIER_1 English(EN) · Xu Han ·

    LC-QAT:通过线性约束向量量化实现 LLM 的数据高效 2 位 QAT

    Quantization-aware training (QAT) is essential for extremely low-bit large language models (LLMs). Current QAT methods are mainly based on scalar quantization (SQ), which enables efficient optimization but suffers from severe performance degradation at 2-bit precision. On the oth…

  6. r/LocalLLaMA TIER_1 (CA) · /u/silenceimpaired ·

    2位元QAT模型发布

    <!-- SC_OFF --><div class="md"><p>So far model releases that take advantage of Quantization a<br /> Aware Training (QAT) have been focused on 4-bit. </p> <p>I’m curious what could be accomplished with a larger MoE model around 120b up to 400b. Obviously the model could not approa…