arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

C-PTQ:用于多模态大语言模型训练后量化的Fisher加权通道敏感性

C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs

Jiameng Li, Han Zhou, Matthew B. Blaschko

arXiv 2607.21076首次发表:更新:

AI 中文总结

研究针对多模态大语言模型训练后量化中异常通道致性能下降问题,提出C-PTQ方法,通过设计Fisher加权目标统一任务特定损失扰动和量化误差,无需辅助模块达最优性能,经实验验证有效。

AI 中文摘要

多模态大语言模型(MLLMs)因内存和计算成本高限制了实际部署,训练后量化(PTQ)技术可解决此问题,但量化模型会因异常通道性能下降。现有PTQ方法利用模态或令牌级指标指导LLM解码器的通道缩放(CWS),但未捕捉通道对任务特定损失的影响。为此提出C-PTQ,统一通道级PTQ方法,通过设计Fisher加权目标将任务敏感性注入缩放过程,无需辅助模块即达最优性能,实验验证了其有效性。

英文摘要

Multimodal large language models (MLLMs) require huge memory and computational costs, which limits their practical deployment. Post-training quantization (PTQ) techniques offer an efficient solution for model compression and inference acceleration. Yet, the quantized model faces performance degradation due to outlier channels, which are highly sensitive to quantization and substantially impair activation fidelity and task accuracy. To protect these salient channels during quantization, existing PTQ methods leverage modality- or token-level metrics to guide channel-wise scaling (CWS) of LLM decoders. However, these orthogonal measurements fail to capture channel-wise impacts on task-specific loss, and the misalignment between importance and scaling factors ultimately leads to suboptimal performance. To address this issue, we propose C-PTQ, a unified channel-wise PTQ method that harmonizes task-specific loss perturbation and quantization error. Motivated by second-order derivatives, we design a Fisher-weighted objective as a tractable Hessian approximation, seamlessly injecting task sensitivity into the scaling process. Notably, we achieve state-of-the-art performance without auxiliary modules like LoRA, thereby maintaining high efficiency. Experiments on Qwen2.5VL, InternVL2 and LLaVA-OV across 8 benchmarks demonstrate our effectiveness in both weight-only and weight-activation settings.

Comments7 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑