arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化ViT的测试时自适应:通过单次前向的量化器对齐重校准

Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration

Hyeongheon Cha, Young D. Kwon, Sung-Ju Lee

arXiv 2610.08358首次发表:更新:

发表机构

KAIST; Samsung AI Center-Cambridge; University of Cambridge(韩国科学技术院; 三星人工智能中心-剑桥; 剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对量化ViT在分布偏移下的脆弱性,提出单次前向、无反向传播的量化器对齐重校准(QuAR)方法,通过重校准激活统计量恢复码本分布,在ImageNet-C上显著提升精度并降低延迟与内存开销。

AI 中文摘要

后训练量化是将视觉变换器(ViT)适配到边缘计算和内存预算的标准途径,然而量化模型在分布偏移下变得特别脆弱。测试时自适应(TTA)无需标签即可应对此类偏移,但大多数现有方法与量化推理的约束条件不太匹配。主流的TTA方法通过反向传播恢复精度,而无反向传播的方法通常仍会因额外的前向传播或参数更新而产生开销,轻量级的特征或logit级方法只能恢复部分损失。在这些方法中,一种量化特有的、会放大精度下降的失败模式并未被直接针对:在偏移下,激活值以不同的方式占据冻结量化器的校准范围,从而扭曲其码本分布。我们提出了量化器对齐重校准(QuAR),一种专为量化ViT设计的单次前向TTA方法,它既不进行反向传播,也不更新任何模型参数。QuAR在冻结量化器的输入处重校准激活值,将测试流的运行逐通道统计量映射回源校准状态。在ImageNet-C上使用ViT-B,QuAR在3位、4位、6位和8位权重/激活精度下,在所有最先进的无反向传播TTA方法中取得了最高的平均精度,在8位时比最强基线高出2.28个百分点,在3位时高出4.00个百分点,同时延迟降低46%,内存开销仅为0.17 MB(峰值推理内存的0.01%)。分析和诊断表明,这一增益源于在这些量化器上减少了逐通道失配,从而恢复了基线方法未改变或进一步扭曲的码本分布。一个固定的配置在连续流、非独立同分布标签偏移、七个分布外测试套件以及另外三个骨干网络上始终保持领先。

英文摘要

Post-training quantization is a standard route to fitting vision transformers (ViTs) into edge compute and memory budgets, yet quantized models become especially brittle under distribution shift. Test-time adaptation (TTA) addresses such shifts without labels, but most existing approaches are poorly aligned with the constraints of quantized inference. Prevailing TTA methods recover accuracy through backpropagation, while backprop-free methods often still incur overhead from extra forward passes or parameter updates, and lightweight feature- or logit-level methods recover only part of the loss. Across these approaches, a quantization-specific failure mode that amplifies the drop is not directly targeted: under shift, activations occupy frozen quantizers' calibrated ranges differently, distorting their code distribution. We propose Quantizer-Aligned Recalibration (QuAR), a single-pass TTA method tailored to quantized ViTs that neither backpropagates nor updates any model parameters. QuAR recalibrates activations at the input to a frozen quantizer, mapping the test stream's running per-channel statistics back toward the source calibration. On ImageNet-C with ViT-B, QuAR achieves the highest mean accuracy among state-of-the-art backprop-free TTA methods at 3-, 4-, 6- and 8-bit weight/activation precision, outperforming the strongest baseline by 2.28 points at 8 bits and 4.00 at 3 bits, with 46% lower latency and a memory overhead of only 0.17 MB (0.01% of peak inference memory). Analysis and diagnostics trace the gain to a reduced per-channel mismatch at these quantizers, which restores the code distribution the baselines leave unchanged or distort further. A single fixed configuration remains ahead across continual streams, non-i.i.d. label shift, seven out-of-distribution suites, and three other backbones.

Comments44 pages, 6 figures. Code at https://github.com/chahh9808/QuAR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑