English(EN)JARQ: Joint Alternating Refinement for Quantization
新研究探索LLM的高级量化技术 · 已追踪10个来源
作者PulseAugur 编辑部·[13 个来源]·
多篇研究论文介绍了量化大型语言模型(LLMs)的新技术,以减少其计算和内存占用。这些方法旨在提高效率,同时不显著牺牲性能。方法包括激活量化的优化旋转、面向量化感知训练的最小范数优化,以及用于部署对齐的量化感知蒸馏框架。其他研究侧重于超低复杂度反量化、重新思考优化损失尺度以及量化的联合交替细化,所有这些都有助于使LLMs更易于访问和更高效。
AI
arXiv:2605.10793v2 Announce Type: replace Abstract: Large language models (LLMs) are costly to deploy due to their large memory footprint and high inference cost. Weight-activation quantization can reduce these costs, but low-bit activation quantization remains difficult because …
arXiv:2610.00738v1 Announce Type: cross Abstract: The Straight-Through Estimator (STE) is a widely used heuristic for Quantization-Aware Training (QAT), but its surrogate gradients can exhibit substantial mismatch with the underlying quantized objective, leading to noisy updates …
arXiv:2609.39223v2 Announce Type: new Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory,…
arXiv:2610.00432v1 Announce Type: cross Abstract: Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponentially large codebooks required by conventional vector quantization. Practical deploy…
arXiv:2610.00983v1 Announce Type: cross Abstract: Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ method…
arXiv:2609.38599v1 Announce Type: new Abstract: Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input corr…
arXiv cs.AI
TIER_1English(EN)·Mehdi Makni, Ryan Lucas, Rahul Mazumder·
arXiv:2609.36120v1 Announce Type: cross Abstract: Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based proc…
arXiv:2609.38169v1 Announce Type: cross Abstract: Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision o…
arXiv cs.LG
TIER_1English(EN)·Jonas von Berg, Massimiliano Datres, Carlo Knei{\ss}l, Gitta Kutyniok·
arXiv:2609.37416v1 Announce Type: new Abstract: Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; the common max-based choice degrades sharply at low bit-widths. We study how sensit…
arXiv:2609.31009v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer …
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with l…
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and …
dev.to — LLM tag
TIER_1English(EN)·Prabhakar Chaudhary·
<h1> REAL-Q: How Dynamic Gradient Descent Fixes the Core Flaw in LLM Quantization </h1> <p>Post-training quantization (PTQ) is one of the most practical tools in the LLM deployment toolkit. Compress a 70B model to 4-bit weights and you can run it on hardware that would otherwise …