PulseAugur
中
实时 09:57:29
English(EN) MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention

MXAttention 框架优化 MXFP4 注意力以用于视频生成

研究人员开发了 MXAttention,这是一个新颖的、无数据的训练后量化框架,旨在优化基于扩散的视频生成模型中的 MXFP4 注意力。该框架解决了数值问题,如裁剪-下溢权衡和归一化错误,这些问题通常会降低生成质量。MXAttention 结合了通用最优缩放 (UOS) 以实现独立于分布的缩放,以及预归一化量化 (PNQ) 以保持归一化精度。实验表明,MXAttention 显著缩小了 MXFP4 和 FP16 格式之间的质量差距,以最小的开销实现了接近 FP16 的生成质量,并且在性能上与 NVFP4 基线相比具有竞争力。 AI

影响 通过优化注意力机制,提高了基于扩散的视频生成模型的效率和质量。

排序理由 详细介绍用于优化 AI 模型性能的新技术框架的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

MXAttention 框架优化 MXFP4 注意力以用于视频生成

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
详细介绍用于优化 AI 模型性能的新技术框架的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
72 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jianlin Yu, Jing Lin, Linghui Kong, Aiyue Chen, Weiyi Sun, Chenyu Zeng, Wangli Lan, Jinxi Li, Zhuo Zheng, Ziyang Yue, Danning Ke, Fei Yi, Tianchi Hu, Yuan Ding, Yiwu Yao, Junsong Wang ·

    MXAttention:MXFP4 Attention 的无数据最优缩放和预归一化量化

    arXiv:2607.24377v1 Announce Type: cross Abstract: The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation qualit…