New research explores advanced quantization techniques for LLMs · 10 sources tracked
ByPulseAugur Editorial·[13 sources]·
Multiple research papers introduce novel techniques for quantizing large language models (LLMs) to reduce their computational and memory footprints. These methods aim to improve efficiency without significantly sacrificing performance. Approaches include optimized rotations for activation quantization, minimal-norm optimization for quantization-aware training, and frameworks for deployment-aligned quantization-aware distillation. Other research focuses on ultra-low-complexity dequantization, rethinking optimization loss scales, and joint alternating refinement for quantization, all contributing to making LLMs more accessible and efficient.
AI
IMPACT
These advancements in LLM quantization aim to significantly reduce inference costs and memory requirements, potentially accelerating the deployment and accessibility of large models across various hardware platforms.
RANK_REASON
Multiple arXiv papers introduce new methods and frameworks for LLM quantization.
arXiv:2605.10793v2 Announce Type: replace Abstract: Large language models (LLMs) are costly to deploy due to their large memory footprint and high inference cost. Weight-activation quantization can reduce these costs, but low-bit activation quantization remains difficult because …
arXiv:2610.00738v1 Announce Type: cross Abstract: The Straight-Through Estimator (STE) is a widely used heuristic for Quantization-Aware Training (QAT), but its surrogate gradients can exhibit substantial mismatch with the underlying quantized objective, leading to noisy updates …
arXiv:2609.39223v2 Announce Type: new Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory,…
arXiv:2610.00432v1 Announce Type: cross Abstract: Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponentially large codebooks required by conventional vector quantization. Practical deploy…
arXiv:2610.00983v1 Announce Type: cross Abstract: Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ method…
arXiv:2609.38599v1 Announce Type: new Abstract: Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input corr…
arXiv cs.AI
TIER_1English(EN)·Mehdi Makni, Ryan Lucas, Rahul Mazumder·
arXiv:2609.36120v1 Announce Type: cross Abstract: Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based proc…
arXiv:2609.38169v1 Announce Type: cross Abstract: Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision o…
arXiv cs.LG
TIER_1English(EN)·Jonas von Berg, Massimiliano Datres, Carlo Knei{\ss}l, Gitta Kutyniok·
arXiv:2609.37416v1 Announce Type: new Abstract: Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; the common max-based choice degrades sharply at low bit-widths. We study how sensit…
arXiv:2609.31009v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer …
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with l…
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and …
dev.to — LLM tag
TIER_1English(EN)·Prabhakar Chaudhary·
<h1> REAL-Q: How Dynamic Gradient Descent Fixes the Core Flaw in LLM Quantization </h1> <p>Post-training quantization (PTQ) is one of the most practical tools in the LLM deployment toolkit. Compress a 70B model to 4-bit weights and you can run it on hardware that would otherwise …