arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GLF-Q:用于视觉Transformer的全局-局部特征量化

GLF-Q: Global-Local Feature-based Quantization for Vision Transformers

Peilin Sun, Guang Liang, Jin Tong, Jianxin Wu

arXiv 2609.34564首次发表:更新:

发表机构

State Key Laboratory of Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University; Zhongguancun Academy(南京大学计算机软件新技术全国重点实验室; 南京大学人工智能学院; 中关村学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低比特后训练量化中ViT精度下降问题,提出GLF-Q框架,通过全局-局部特征对齐与离线Hadamard变换,在3比特量化下显著超越现有方法,并具备鲁棒性与加速优势。

AI 中文摘要

后训练量化(PTQ)无需重新训练即可高效压缩视觉Transformer(ViT),但在低比特宽度下会遭受严重的精度下降。现有的基于优化的PTQ方法通过软逻辑或二阶Hessian代理来指导块重建。逻辑监督容易在有限的校准数据上过拟合,而Hessian近似则会产生结构截断误差。为解决这些局限性,我们提出了GLF-Q,一种由全局-局部特征对齐引导的新型PTQ框架。GLF-Q将量化块的输出通过下游全精度层传播,在局部输出正则化下对齐倒数第二层表示,从而在不显式近似Hessian或使用泰勒展开的情况下提供下游特征监督。此外,引入了离线Hadamard变换,且运行时开销为零,以分散跨通道的激活异常值,有效收缩动态范围并减少量化误差。同时,通过直通估计器(STE)优化该损失可实现快速收敛,绕过了诸如AdaRound之类的连续松弛舍入公式。在代表性ViT架构上的大量实验表明,在图像分类的3比特量化下,采用标准均匀量化器的GLF-Q显著优于最先进的方法。此外,GLF-Q表现出强大的域外校准鲁棒性,并在8比特GPU部署下实现了加速。

英文摘要

Post-training quantization (PTQ) efficiently compresses Vision Transformers (ViTs) without retraining, yet suffers severe accuracy degradation at low bit-widths. Existing optimization-based PTQ methods guide block reconstruction via either soft logits or second-order Hessian proxies. Logit supervision is prone to overfitting on limited calibration data, while Hessian approximations incur structural truncation errors. To address these limitations, we propose \textbf{GLF-Q}, a novel PTQ framework guided by Global-Local Feature alignment. GLF-Q propagates quantized block outputs through downstream full-precision layers to align penultimate-layer representations under local output regularization, providing downstream feature supervision without explicitly approximating the Hessian or using a Taylor expansion. Furthermore, offline Hadamard transformations are introduced with zero runtime overhead to disperse activation outliers across channels, effectively contracting dynamic ranges and reducing quantization errors. Meanwhile, optimizing this loss via a Straight-Through Estimator (STE) achieves rapid convergence, bypassing continuous relaxation rounding formulations such as AdaRound. Extensive experiments across representative ViT architectures demonstrate that GLF-Q with standard uniform quantizers substantially outperforms state-of-the-art methods under 3-bit quantization on image classification. In addition, GLF-Q exhibits strong out-of-domain calibration robustness and achieves speedups under 8-bit GPU deployment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑