PulseAugur
实时 09:42:21
English(EN) Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models

新方法显著提升大语言模型低比特量化性能

研究人员开发了一种新颖的大语言模型极低比特量化方法,解决了跨层累积误差导致性能下降的问题。他们的方法通过联合优化所有层的离散码和尺度,并结合跨层误差补偿和有限样本特征统计匹配。该技术显著提高了性能,在 Qwen2.5-1.5B 模型上实现了 1.125 比特权重的困惑度比 9.56,大幅优于现有方法。 AI

影响 能够更有效地在资源受限的硬件上部署大语言模型。

排序理由 关于大语言模型量化新技术的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.NE (Neural & Evolutionary) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新方法显著提升大语言模型低比特量化性能

报道来源 [1]

  1. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · Ryona Noda ·

    大型语言模型的极端低比特量化中的跨层误差补偿与有限样本特征统计匹配

    Layer-wise post-training quantization of large language models minimizes each layer's reconstruction error in isolation, allowing quantization errors to accumulate across depth and causing severe degradation in extreme low-bit regimes. We formulate quantization as a joint optimiz…