MXAttention:用于MXFP4注意力的数据无关最优缩放和预归一化量化
MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention
浏览论文内容
中文总结 AI 辅助
针对基于扩散的视频生成模型中注意力二次成本的瓶颈,提出MXAttention框架,通过引入通用最优缩放和预归一化量化组件,缩小了MXFP4与FP16的成像质量差距,提升了帧级相似度,且性能与强大基线竞争。
中文摘要 AI 辅助
注意力的二次成本是基于扩散的视频生成模型的主要瓶颈。MXFP4注意力为高效推理提供了一条有前景的途径,但直接的MXFP4量化常因两个数值问题而降低生成质量。我们提出MXAttention,一种用于MXFP4注意力的数据无关的训练后量化框架。它引入通用最优缩放和预归一化量化两个组件。实验表明,MXAttention缩小了OCP MXFP4与FP16之间至少95%的VBench成像质量差距,显著提高帧级相似度,且在所有VBench指标上绝对降级小于0.01,还实现了与基于NVFP4的强大基线竞争的性能。
英文摘要
The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path toward efficient inference, but direct MXFP4 quantization often degrades generation quality due to two numerical issues: the clipping-underflow trade-off from power-of-two scaling and the row-wise normalization error introduced in the softmax loop. We propose MXAttention, a data-free post-training quantization framework for MXFP4 attention. MXAttention introduces two components: Universal Optimal Scaling (UOS), which exploits the periodic structure of power-of-two microscaling to derive a distribution-independent optimal scaling boundary Qmax=7.25 without calibration or search, and Pre-Normalization Quantization (PNQ), which quantizes unnormalized softmax exponentials before row-wise summation to preserve normalization by construction. Experiments on Wan2.2 and HunyuanVideo show that MXAttention closes at least 95% of the VBench Imaging Quality gap between OCP MXFP4 and FP16, substantially improves frame-level similarity, and preserves FP16-level generation quality with less than 0.01 absolute degradation on all reported VBench metrics. MXAttention also achieves performance competitive with strong NVFP4-based baselines with negligible overhead when fused into the attention pipeline. The implementation is publicly available in MindIE-SD.
发表机构
- Huawei Technologies Co., Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。