发表机构
Shanghai Jiao Tong University; The Hong Kong Polytechnic University(上海交通大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SPHQuant提出无旋转球面权重量化框架,通过分解权重向量隔离异常值,提升VLM极低比特量化精度,并实现30.3%的解码吞吐提升。
AI 中文摘要
最近的基础模型正朝着原生多模态视觉-语言模型(VLM)的方向发展,使得VLM成为下一代基础模型的核心形式。然而,其庞大的语言骨干网络由于高内存占用和受内存限制的自回归解码,使得边缘部署变得困难。仅权重的训练后量化是一种实用的解决方案,但将VLM推向极低比特宽度仍然具有挑战性:现有的无旋转方法在2-3比特下受异常值影响,而基于旋转的方法虽提高了精度,却增加了额外的运行时开销。我们提出了SPHQuant,一种用于VLM的无旋转球面权重量化框架。SPHQuant不是直接在笛卡尔坐标中量化权重,而是将每个8维权重向量分解为坐标符号、半径和正单位方向。这种表示将异常值幅度隔离到半径中,同时保持方向有界且统计规律。基于这一见解,SPHQuant为半径分配额外精度以缓解异常值引起的精度下降。它进一步使用紧凑的正方向码本,并通过角度参数化微调码本条目以保持单位球约束。我们还设计了一个硬件友好的GEMV内核,使方向码本足够小以进行共享内存查找,并高效地打包径向比特。实验表明,SPHQuant在匹配最先进的极低比特量化方法性能的同时,在RTX A6000上将解码吞吐量比QTIP提高了30.3%。代码将在此https URL中发布。
英文摘要
Recent foundation models are moving toward native multimodal Vision-Language Models (VLMs), making VLMs a central form of next-generation foundation models. However, their large language backbones make edge deployment difficult due to high memory footprint and memory-bound autoregressive decoding. Weight-only post-training quantization is a practical solution, but pushing VLMs to extreme low bit-widths remains challenging: existing rotation-free methods suffer from outliers at 2-3 bits, while rotation-based methods improve accuracy at the cost of additional runtime overhead. We propose SPHQuant, a rotation-free spherical weight-only quantization framework for VLMs. Instead of quantizing weights directly in Cartesian coordinates, SPHQuant decomposes each 8D weight vector into coordinate signs, radius, and a positive unit direction. This representation isolates outlier magnitude into the radius while keeping directions bounded and statistically regular. Based on this insight, SPHQuant allocates extra precision to the radius to mitigate accuracy degradation induced by outliers. It further uses a compact positive-direction codebook and fine-tunes codebook entries through angular parameterization to preserve the unit-sphere constraint. We also design a hardware-friendly GEMV kernel that keeps the direction codebook small enough for shared-memory lookup and packs radial bits efficiently. Experiments show that SPHQuant matches the performance of state-of-the-art extreme low-bit quantization methods while improving decode throughput over QTIP by 30.3% on RTX A6000. Code will be released in https://github.com/Pushazf/SPHQuant.
Comments16 pages, 5 figures, including appendix. Code will be released at https://github.com/Pushazf/SPHQuant