arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RGSQ:面向大型视觉-语言模型的黎曼几何敏感量化

RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models

Zhiping Wu, Dongdong Ren, Yangchengyu Zhou, Zhengjie Zhang, Wenbin Li, Hongbing Pan, Yang Gao

arXiv 2609.25492首次发表:更新:

发表机构

Nanjing University; Geely Automobile Research Institute (Ningbo) Co., Ltd.(南京大学; 吉利汽车研究院(宁波)有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RGSQ,利用Fisher-Riemannian度量与几何旋转及白化变换,在极低比特下提升VLM量化精度,优于现有基线。

AI 中文摘要

大型视觉-语言模型(VLM)在严格的内存和延迟约束下,可通过训练后量化(PTQ)高效部署。然而,大多数PTQ方法是为单模态大型语言模型(LLM)设计的。这些方法在欧几里得假设下将量化误差视为各向同性扰动,这为VLM中对量化最敏感的方向提供了较弱的指导。因此,直接采用单模态PTQ方法或仅使用模态特定缩放,往往会导致低比特设置下比特宽度分布不均和性能不一致。为解决这些挑战,我们提出了黎曼几何敏感量化(RGSQ),该方法将量化表述为在统一Fisher-Riemannian度量下的重构问题。RGSQ通过由模态分区经验Fisher因子构建的黎曼流形映射,识别模态特定的敏感方向,并将其融合为模态感知的Kronecker结构度量。然后,我们应用几何对齐旋转来重新定向局部切空间框架,将低比特扰动引导至损失不敏感的轴。最后,我们应用白化变换,将黎曼目标映射为等效的欧几里得形式,使标准单模态PTQ方法能够在原始假设下评估多模态量化误差。在广泛且多样化的主流VLM基准测试中,RGSQ在极低比特设置(W2A8和W3A8)下实现了最高准确性和稳定性。它比VLM感知基线(如MBQ和MQuant)最多高出5.9%,并比单模态改进最多高出8.6%。

英文摘要

Large vision-language models (VLMs) can be efficiently deployed under stringent memory and latency constraints through post training quantization (PTQ). However, most PTQ methods are designed for unimodal large language models (LLMs). These methods treat quantization errors as isotropic perturbations under the Euclidean assumption, which provides weak guidance on directions most sensitive to quantization in VLMs. Consequently, directly adapting unimodal PTQ approaches or solely employing modality-specific scaling often leads to uneven bit-width distribution and inconsistent performance in low-bit settings. To address these challenges, we propose Riemannian Geometry-Sensitive Quantization (RGSQ), which formulates quantization as a reconstruction problem under a unified Fisher-Riemannian metric. RGSQ identifies modality-specific sensitive directions via Riemannian manifold mappings built from modality-partitioned empirical Fisher factors and fused into a modality-aware Kronecker-structured metric. We then apply geometry-aligned rotations to reorient the local tangent frame, steering low-bit perturbations toward loss-insensitive axes. Finally, we apply a whitening transformation that maps the Riemannian objective to an equivalent Euclidean form, enabling standard unimodal PTQ methods to evaluate multimodal quantization error under their original assumptions. Across an extensive and diverse set of mainstream VLM benchmarks, RGSQ achieves the highest accuracy and stability under extremely low-bit settings (W2A8 and W3A8). It outperforms VLM-aware baselines, such as MBQ and MQuant, by up to 5.9% and surpasses single-modality improvements by up to 8.6%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑