English(EN)Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression
新的LLM压缩技术带来更小、更精确的模型
作者PulseAugur 编辑部·[12 个来源]·
研究人员开发了新的方法来压缩大型语言模型(LLM),同时保持甚至提高其性能。一种方法是量化感知修复(QAH),它直接从原始未压缩模型中蒸馏出一个压缩的4位模型,从而得到一个更小、更精确、更高效的模型,在多个基准测试中优于其全精度对应模型。其他研究探索了激活加权种子残差编码(AWSRC)等技术来修复量化误差,以及一个名为“压缩三位一体”的统一框架,该框架联合应用稀疏性、量化和低秩近似来实现高效的LLM部署。此外,还创建了一个标准化的评估平台LowRankArena,以促进LLM压缩方法的可复现比较。
AI
arXiv:2608.26389v1 Announce Type: new Abstract: SVD-based low-rank compression has become a fast-growing direction for reducing the memory and computational cost of large language models (LLMs). However, meaningful comparison across existing studies remains difficult as prior eva…
arXiv cs.AI
TIER_1English(EN)·Evangelos Georganas, Dhiraj Kalamkar, Alexander Heinecke, Pradeep Dubey·
arXiv:2508.06753v3 Announce Type: replace Abstract: The advent of ultra-low-bit LLM models, approaching the perplexity and task accuracy of their full precision counterparts, is ushering in a new era of LLM inference. While these advances promise models that are cost-effective re…
arXiv cs.LG
TIER_1English(EN)·Tanzila Rahman, Mehran Taghian Jazi, Yunke Peng, Zhuang Ma, Anandharaju Durai Raju, Yao Wang, Xing Huang, Hei Yi Mak, Shadan Golestan, Hoang Le, Yonghan Dong, Wei Guo, Yaoyuan Wang·
arXiv:2608.26581v1 Announce Type: new Abstract: Low-bit quantization offers a promising avenue for reducing the computational and memory demands of Multimodal Large Language Models (MLLMs). Recent hardware support for low-precision formats, ranging from MXFP8 to ultra-low-bit for…
arXiv:2608.25188v1 Announce Type: new Abstract: Most competitive 4-bit LLM research pipelines now open the same way: apply a linear, function-preserving transform (rotation, scaling, permutation, non-orthogonal affine) so the outlier mass sits more favorably against the group sca…
arXiv:2608.24070v1 Announce Type: new Abstract: Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolatio…
Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. Th…
arXiv cs.CL
TIER_1English(EN)·Zehao Liu, Chuangchuang Fang, Yang Ren·
arXiv:2608.23144v1 Announce Type: cross Abstract: Low-bit weight quantization saves storage but leaves errors that degrade language-model quality. We introduce Activation-Weighted Seeded Residual Coding (AWSRC), a compact repair codec for an existing quantization backbone. Given …
arXiv cs.AI
TIER_1English(EN)·Bakbergen Ryskulov, Iker Garc\'ia-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Rom\'an Or\'us·
arXiv:2608.20953v1 Announce Type: cross Abstract: Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding,…
Quantization-aware healing recovers compressed 4-bit language models faster and more stably than quantization-aware training by distilling directly from the original uncompressed model.
<p>A new technique, Quantization-Aware Healing (QAH), has been introduced, enabling 4-bit quantized models to surprisingly surpass the performance of their full-precision counterparts. This breakthrough directly addresses the challenge of deploying large language models efficient…