arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24377cs.LGcs.AIcs.CV

MXAttention:用于MXFP4注意力的数据无关最优缩放和预归一化量化

MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention

Jianlin Yu, Jing Lin, Linghui Kong, Aiyue Chen, Weiyi Sun, Chenyu Zeng, Wangli Lan, Jinxi Li, Zhuo Zheng, Ziyang Yue, Danning Ke, Fei Yi, Tianchi Hu, Yuan Ding,… 展开作者

Jianlin Yu, Jing Lin, Linghui Kong, Aiyue Chen, Weiyi Sun, Chenyu Zeng, Wangli Lan, Jinxi Li, Zhuo Zheng, Ziyang Yue, Danning Ke, Fei Yi, Tianchi Hu, Yuan Ding, Yiwu Yao, Junsong Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对基于扩散的视频生成模型中注意力二次成本的瓶颈,提出MXAttention框架,通过引入通用最优缩放和预归一化量化组件,缩小了MXFP4与FP16的成像质量差距,提升了帧级相似度,且性能与强大基线竞争。

中文摘要 AI 辅助

注意力的二次成本是基于扩散的视频生成模型的主要瓶颈。MXFP4注意力为高效推理提供了一条有前景的途径,但直接的MXFP4量化常因两个数值问题而降低生成质量。我们提出MXAttention,一种用于MXFP4注意力的数据无关的训练后量化框架。它引入通用最优缩放和预归一化量化两个组件。实验表明,MXAttention缩小了OCP MXFP4与FP16之间至少95%的VBench成像质量差距,显著提高帧级相似度,且在所有VBench指标上绝对降级小于0.01,还实现了与基于NVFP4的强大基线竞争的性能。

英文摘要

The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical issues: the clipping-underflow trade-off from power-of-two scaling and the row-wise normalization error introduced in the softmax loop. We propose MXAttention, a data-free post-training quantization framework for MXFP4 attention. MXAttention introduces two components: Universal Optimal Scaling (UOS), which exploits the periodic structure of power-of-two microscaling to derive a distribution-independent optimal scaling boundary Qmax=7.25 without calibration or search, and Pre-Normalization Quantization (PNQ), which quantizes unnormalized softmax exponentials before row-wise summation to preserve normalization by construction. Experiments on Wan2.2 and HunyuanVideo show that MXAttention closes at least 95% of the VBench Imaging Quality gap between OCP MXFP4 and FP16, substantially improves frame-level similarity, and preserves FP16-level generation quality with less than 0.01 absolute degradation on all reported VBench metrics. MXAttention also achieves performance competitive with strong NVFP4-based baselines with negligible overhead when fused into the attention pipeline. The implementation is publicly available in MindIE-SD.

发表机构

  • Huawei Technologies Co., Ltd.(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑