PulseAugur
实时 06:09:46
English(EN) RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

RotaryQuant 使 120B MoE 模型能在消费级硬件上运行

研究人员开发了 RotaryQuant,一个旨在使大型专家混合(MoE)语言模型能在消费级硬件上运行的新型压缩系统。该系统采用三轴压缩策略,包括混合精度权重量化、LRU 专家卸载以及用于 KV 缓存压缩的 IsoQuant。这种方法允许 Nemotron-H 120B 等模型在 32GB 内存预算内运行,同时保持接近零的困惑度下降和高检索准确率。 AI

影响 使大型 MoE 模型能在消费级硬件上运行,可能使先进的 AI 能力的获取更加普及。

排序理由 该集群描述了一种新颖的压缩技术,该技术在 arXiv 论文中提出,用于将大型语言模型适配到消费级硬件。

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

RotaryQuant 使 120B MoE 模型能在消费级硬件上运行

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一种新颖的压缩技术,该技术在 arXiv 论文中提出,用于将大型语言模型适配到消费级硬件。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
infra, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
17 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [2]

  1. arXiv cs.LG TIER_1 English(EN) · Anthony. Lui, Mohamed. Elsaied, N. P. Savani ·

    RotaryQuant:通过融合压缩空间注意力将 120B MoE 模型安装到消费级硬件上

    arXiv:2608.08081v1 Announce Type: cross Abstract: Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows li…

  2. arXiv cs.NE (Neural & Evolutionary) TIER_1 English(EN) · N. P. Savani ·

    RotaryQuant:通过融合压缩空间注意力将 120B MoE 模型安装到消费级硬件上

    Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayer…