PulseAugur
中
实时 04:29:28
English(EN) Tetra: Serving Leech-Lattice Quantized LLMs at 2.7 Bits per Parameter

Tetra LLM 量化实现每参数 2.7 比特,提升推理速度

研究人员开发了 Tetra,一种将大语言模型(LLM)量化至平均每参数 2.7 比特的新颖方法,显著降低了内存需求。该技术基于 leech-lattice 量化,通过优化码本结构并采用专门的内核来最小化 GPU 内存加载的数据,从而实现这一目标。使用 Tetra 量化的模型,如 Qwen3 变体,在 MMLU 和 GSM8K 等基准测试中表现出与更高比特量化方法相当的性能,仅有轻微下降,同时在推理速度方面显示出显著提升。 AI

影响 通过显著减小内存占用和提高推理速度,实现了更高效的大语言模型部署。

排序理由 详细介绍大语言模型量化新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Tetra LLM 量化实现每参数 2.7 比特,提升推理速度

报道来源 [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    Tetra:以每参数2.7比特服务蛭状格量化大语言模型

    Leech-lattice quantization gives good quality at two bits per weight, but its codebooks hold more than 10^14 points, too many for a lookup table. Our earlier kernel expanded the codes at load time and read 4.804 bits per weight from GPU memory for 2 bits of code. We present Tetra…