PulseAugur
中
实时 18:20:14
English(EN) QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides

新的QUADS技术稳定MoE LLM的NVFP4强化学习

研究人员开发了一种名为QUADS的新技术,通过使用NVFP4低精度格式来稳定混合专家(MoE)大语言模型的强化学习(RL)。他们发现,激活误差而非权重误差是MoE模型FP4 RL不稳定的主要原因。QUADS通过在训练器端实现不对称量化感知训练,并在回滚端实现残差激活补偿来解决此问题,达到了BF16级别的准确率并提高了吞吐量。 AI

影响 这项研究通过允许使用低精度格式而不牺牲准确性,有望实现更高效的大语言模型训练。

排序理由 这是一篇研究论文,详细介绍了一种稳定MoE LLM中强化学习的新技术。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的QUADS技术稳定MoE LLM的NVFP4强化学习

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是一篇研究论文,详细介绍了一种稳定MoE LLM中强化学习的新技术。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
80 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Zhengyang Zhuge, Hao Yu, Xin Wang, Zheng Li, Yizhong Cao, Dayiheng Liu, Jianwei Zhang ·

    QUADS:通过双边量化误差对齐稳定MoE的NVFP4强化学习

    arXiv:2607.15810v1 Announce Type: new Abstract: Rollout generation is a major bottleneck in Reinforcement Learning (RL) for Mixture-of-Experts (MoE) Large Language Models, motivating low-precision rollout acceleration such as FP8. As an emerging low-precision format, NVFP4 combin…