arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34884cs.CV

SubRot:用于VLM旋转量化的符号梯度子空间校准

SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization

Zhenhao Shang, Haizhao Jing, Haokui Zhang, Guoting Wei, Rong Xiao, Jianqing Gao, Peng Wang

首次发表
浏览论文内容

中文总结 AI 辅助

提出SubRot方法,利用符号梯度子空间校准VLM旋转量化,通过特征分解和泰勒展开引导误差方向,在多个基准上提升量化模型性能。

中文摘要 AI 辅助

后训练量化降低了视觉语言模型(VLM)的部署成本,但在低比特宽度下保持多模态能力仍然具有挑战性。现有方法依赖于模态级或词元级的梯度统计,这些统计容易受到视觉到文本词元比例和视觉信息位置的跨样本变化的影响,从而限制了统计稳定性。此外,通过绝对值求平均进行的过度粗糙聚合丢弃了梯度符号和通道间差异,限制了对模态特定敏感性的分离。相比之下,通道空间为跨样本提供了共享的坐标系,使其成为捕获稳定任务敏感结构的更自然基础。因此,我们提出了SubRot,一种用于VLM旋转量化的符号梯度子空间校准方法。通过对激活梯度的经验Fisher矩阵进行特征分解,SubRot识别出一个敏感通道子空间,该子空间具有三个特性:跨样本稳定性、清晰的敏感性分离以及沿某些方向对自回归损失的一致符号效应。在局部泰勒展开的指导下,SubRot将沿符号稳定方向的符号一阶引导与沿其余敏感方向的二阶约束相结合,同时保留MSE以维持整体重建质量。该目标将量化误差引导至损失减少方向,同时控制其幅度。在五个基准上的五个VLM上的实验显示,在W4A6和W4A4设置下,与FlatQuant相比,平均分数持续提升,在LLaVA-NeXT-7B上达到1.4个百分点。在W4A4下,所有评估模型的平均精度下降相对于FP16保持在1.4个百分点以内,而LLaVA-v1.5-13B的平均分数超过其FP16平均分数0.4个百分点。

英文摘要

Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios and the positions of visual information, limiting statistical stability. Moreover, overly coarse aggregation through absolute values and averaging discards gradient signs and channel-wise differences, limiting the separation of modality-specific sensitivities. In contrast, the channel space provides a shared coordinate system across samples, making it a more natural basis for capturing stable task-sensitive structures. We therefore propose SubRot, a signed gradient subspace calibration method for VLM rotation quantization. Through eigendecomposition of the empirical Fisher matrix of activation gradients, SubRot identifies a sensitive channel subspace with three properties: cross-sample stability, clear sensitivity separation, and consistent signed effects on the autoregressive loss along certain directions. Guided by a local Taylor expansion, SubRot combines signed first-order guidance along sign-stable directions with second-order constraints along the remaining sensitive directions, while retaining MSE for overall reconstruction quality. This objective steers quantization errors toward loss-decreasing directions while controlling their magnitude. Experiments on five VLMs across five benchmarks show consistent average-score improvements over FlatQuant under W4A6 and W4A4, reaching 1.4 percentage points on LLaVA-NeXT-7B. Under W4A4, average accuracy degradation from FP16 remains within 1.4 percentage points across all evaluated models, while LLaVA-v1.5-13B exceeds its FP16 average score by 0.4 percentage points.

发表机构

  • Northwest Polytechnical University(西北工业大学)
  • Nanjing University of Science and Technology(南京理工大学)
  • Intellifusion(云天励飞)
  • iFLYTEK CO., LTD(科大讯飞股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑