发表机构
Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对多模态大语言模型超低比特量化的性能损失问题,提出Residual Fallback Quantization框架,可有效恢复MXFP4、HiF4量化下的性能,缩小与BF16基线的差距。
AI 中文摘要
低比特量化为降低多模态大语言模型(Multimodal Large Language Models, MLLMs)的计算与内存需求提供了极具前景的途径。近期对从MXFP8到MXFP4、HiF4等超低比特格式的低精度格式的硬件支持,推动了高效MLLM训练与部署的研究。本研究对覆盖视频生成与推理任务的代表性MLLMs中的上述量化方案开展系统研究,分析显示MXFP8可实现近无损性能,而激进的4比特量化会导致显著性能下降。通过大量 ablation( ablation 指消融实验,即控制变量的实验方法),我们确定激活量化是该性能损失的主要来源,其贡献远大于权重量化。基于此观察,我们提出Residual Fallback Quantization(RFQ,即残差 fallback 量化),这是一种轻量型激活重建框架,通过辅助量化残差路径补充主超低比特激活表示,显式建模并补偿量化误差,在保留超低比特计算效率优势的同时提升激活保真度。RFQ无需架构修改,仅产生可忽略的计算开销。在Wan2.2和Qwen3-VL上的大量实验表明,RFQ可一致恢复MXFP4与HiF4量化下损失的大部分性能,显著缩小了生成与推理基准上与BF16基线的差距。本研究明确激活量化是超低比特MLLMs的主要瓶颈,并凸显基于残差的激活重建是实现鲁棒4比特部署的有效且实用策略。
英文摘要
Low-bit quantization offers a promising avenue for reducing the computational and memory demands of Multimodal Large Language Models (MLLMs). Recent hardware support for low-precision formats, ranging from MXFP8 to ultra-low-bit formats such as MXFP4 and HiF4, has accelerated research into efficient MLLM training and deployment. In this work, we present a systematic study of these quantization schemes in representative MLLMs that span both video generation and reasoning tasks. Our analysis shows that MXFP8 achieves near-lossless performance, whereas aggressive 4-bit quantization leads to significant degradation. Through extensive ablations, we identify activation quantization as the primary source of this performance loss, contributing substantially more than weight quantization. Motivated by this observation, we propose Residual Fallback Quantization (RFQ), a lightweight activation reconstruction framework that supplements the primary ulta-low-bit activation representation with an auxiliary quantized residual pathway. By explicitly modeling and compensating for quantization errors, RFQ improves activation fidelity while preserving the efficiency advantages of ultra-low-bit computation. RFQ requires no architectural modifications and incurs negligible computational overhead. Extensive experiments on Wan2.2 and Qwen3-VL demonstrate that RFQ consistently recovers a substantial portion of the performance lost under the quantization of MXFP4 and HiF4, significantly narrowing the gap to BF16 baselines across both generation and 4 reasoning benchmarks. Our findings establish activation quantization as the dominant bottleneck in ultra-low-bit MLLMs and highlight residual-based activation reconstruction as an effective and practical strategy for robust 4-bit deployment.
Comments14 Pages, 5 figures, 5 tables