arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.16590cs.CLcs.AI

BATQuant: 通过可学习的分块优化实现抗异常的MXFP4量化

BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization

  • Huawei Technologies University of Science(华为技术大学科学)

机构由 AI 辅助整理,请以论文原文为准。

Ji-Fu Li, Manyi Zhang, Xiaobo Xia, Han Bao, Haoli Bai, Zhenhua Dong, Xianzhi Yu

更新

AI总结:

本文提出BATQuant方法,通过可学习的分块优化解决MXFP4量化中异常传播问题,实现高性能量化方案。

AI中文摘要:

微缩浮点(MXFP)格式已成为部署多模态大语言模型(MLLMs)和大语言模型(LLMs)在现代加速器架构中的有希望标准。然而,现有的训练后量化(PTQ)方法,特别是针对整数格式设计的旋转方法,在应用于MXFP4时会遭受严重的性能崩溃。最近的研究将这种失败归因于根本的格式不匹配:全局正交旋转会无意中将异常能量跨量化块转移,产生新的异常体,破坏局部分块缩放,同时往往形成双峰激活分布,导致量化范围的低效利用。为了解决这些问题,我们提出了BATQuant(分块仿射变换),该方法限制变换以与MXFP粒度对齐,以防止跨块异常传播,同时放松正交约束以优化分布形状。为了确保参数效率,我们引入了全局和私有克罗内克(GPK)分解,以有效减少存储和运行时开销,并结合分块可学习截断以抑制残留异常。在MLLMs和LLMs上的大量实验表明,BATQuant在激进的W4A4KV16配置下建立了新的最先进结果,在多模态基准上恢复了高达96.43%的全精度性能,并在各种任务中明显优于现有方法。

英文摘要:

Microscaling floating-point (MXFP) formats have emerged as a promising standard for deploying Multi-modal Large Language Models (MLLMs) and Large Language Models (LLMs) on modern accelerator architectures. However, existing Post-Training Quantization (PTQ) methods, particularly rotation-based techniques designed for integer formats, suffer from severe performance collapse when applied to MXFP4. Recent studies attribute this failure to a fundamental format mismatch: global orthogonal rotations inadvertently transfer outlier energy across quantization blocks, inducing new outliers that disrupt local block-wise scaling, while often creating bimodal activation distributions that underutilize the limited quantization range. To address these issues, we propose BATQuant (Block-wise Affine Transformation), which restricts transformations to align with MXFP granularity to prevent cross-block outlier propagation, while relaxing orthogonality constraints to optimize distribution shaping. To ensure parameter efficiency, we introduce Global and Private Kronecker (GPK) decomposition to effectively reduces storage and runtime overhead and incorporate Block-wise Learnable Clipping to suppress residual outliers. Extensive experiments on both MLLMs and LLMs demonstrate that BATQuant establishes new state-of-the-art results under aggressive W4A4KV16 configurations, recovering up to 96.43% of full-precision performance on multimodal benchmarks and clearly outperforming existing methods across diverse tasks.

补充信息

↑