PulseAugur
中
实时 05:00:44
English(EN) JARQ: Joint Alternating Refinement for Quantization

新研究探索LLM的高级量化技术 · 已追踪10个来源

多篇研究论文介绍了量化大型语言模型(LLMs)的新技术,以减少其计算和内存占用。这些方法旨在提高效率,同时不显著牺牲性能。方法包括激活量化的优化旋转、面向量化感知训练的最小范数优化,以及用于部署对齐的量化感知蒸馏框架。其他研究侧重于超低复杂度反量化、重新思考优化损失尺度以及量化的联合交替细化,所有这些都有助于使LLMs更易于访问和更高效。 AI

影响 LLM量化的这些进展旨在显著降低推理成本和内存需求,有可能加速大型模型在各种硬件平台上的部署和可访问性。

排序理由 多篇arXiv论文介绍了LLM量化的新方法和框架。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 13 个来源。 我们如何撰写摘要 →

新研究探索LLM的高级量化技术 · 已追踪10个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇arXiv论文介绍了LLM量化的新方法和框架。
Source corroboration
13 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
16 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [13]

  1. arXiv cs.LG TIER_1 English(EN) · Chayne Thrash, Ali Abbasi, Soheil Kolouri ·

    ConQuR: 通过优化的旋转实现角落对齐激活量化,用于大型语言模型

    arXiv:2605.10793v2 Announce Type: replace Abstract: Large language models (LLMs) are costly to deploy due to their large memory footprint and high inference cost. Weight-activation quantization can reduce these costs, but low-bit activation quantization remains difficult because …

  2. arXiv cs.LG TIER_1 English(EN) · Don Li ·

    Q-MINO:一种用于量化感知训练的最小范数方法

    arXiv:2610.00738v1 Announce Type: cross Abstract: The Straight-Through Estimator (STE) is a widely used heuristic for Quantization-Aware Training (QAT), but its surrogate gradients can exhibit substantial mismatch with the underlying quantized objective, leading to noisy updates …

  3. arXiv cs.LG TIER_1 English(EN) · Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun ·

    QATFactory:用于 LLM 量化感知训练和蒸馏的多功能、面向部署的框架

    arXiv:2609.39223v2 Announce Type: new Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory,…

  4. arXiv cs.AI TIER_1 English(EN) · Xiaofan Que, Nir Elkayam, Spandan Pyakurel, Shuokai Pan, Dibakar Gope ·

    XOR-Trellis:超低复杂度量化和感知曲率的无哈达玛变换大语言模型量化

    arXiv:2610.00432v1 Announce Type: cross Abstract: Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponentially large codebooks required by conventional vector quantization. Practical deploy…

  5. arXiv cs.AI TIER_1 English(EN) · Chao Li, Shigeng Wang, Anbang Yao ·

    魔鬼藏在重构损失尺度中:重新思考LLM量化中的优化

    arXiv:2610.00983v1 Announce Type: cross Abstract: Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ method…

  6. arXiv cs.LG TIER_1 English(EN) · Xinyu Wang, Sicheng Lyu, Xiao-Wen Chang ·

    JARQ:联合交替量化精炼

    arXiv:2609.38599v1 Announce Type: new Abstract: Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input corr…

  7. arXiv cs.AI TIER_1 English(EN) · Mehdi Makni, Ryan Lucas, Rahul Mazumder ·

    ThinQuant:LLM权重和激活量化的可扩展旋转学习

    arXiv:2609.36120v1 Announce Type: cross Abstract: Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based proc…

  8. arXiv cs.AI TIER_1 English(EN) · Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu, Zhenan Sun, Ying Wei ·

    STEPQuant:Delta-Rule循环状态量化中的错误何时何地至关重要

    arXiv:2609.38169v1 Announce Type: cross Abstract: Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision o…

  9. arXiv cs.LG TIER_1 English(EN) · Jonas von Berg, Massimiliano Datres, Carlo Knei{\ss}l, Gitta Kutyniok ·

    低比特训练后量化中的尺度敏感性:量化误差景观的曲率

    arXiv:2609.37416v1 Announce Type: new Abstract: Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; the common max-based choice degrades sharply at low bit-widths. We study how sensit…

  10. arXiv cs.AI TIER_1 English(EN) · Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang, Xiangsheng Zhou ·

    G$^2$PTQ: 通过广义梯度补偿改进 LLM 训练后量化

    arXiv:2609.31009v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer …

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    G^2PTQ:通过广义梯度补偿改进LLM训练后量化

    Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with l…

  12. Hugging Face Daily Papers TIER_1 Italiano(IT) ·

    Disaggregated Quantization: Specializing LLM Prefill and Decode

    Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and …

  13. dev.to — LLM tag TIER_1 English(EN) · Prabhakar Chaudhary ·

    REAL-Q:动态梯度下降如何修复LLM量化的核心缺陷

    <h1> REAL-Q: How Dynamic Gradient Descent Fixes the Core Flaw in LLM Quantization </h1> <p>Post-training quantization (PTQ) is one of the most practical tools in the LLM deployment toolkit. Compress a 70B model to 4-bit weights and you can run it on hardware that would otherwise …