发表机构
Khulna University of Engineering & Technology (KUET)(库尔纳工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有PTQ方法精度分配低效问题,提出MixFrag框架,通过KL散度估计量化脆弱性并将位分配建模为MCKP,在ImageNet、COCO任务上实现最优混合精度PTQ性能。
AI 中文摘要
后训练量化(PTQ)已成为将视觉Transformer(ViTs)部署到资源受限设备上的有效解决方案。然而,现有的PTQ方法通常在Transformer各组件中采用统一的位宽,忽略了它们对量化的异构敏感性,导致精度分配效率低下。本文提出了MixFrag——一种面向视觉Transformer的脆弱性引导混合精度PTQ框架。MixFrag首先通过使用小型校准集测量全精度输出分布与孤立量化输出分布之间的Kullback-Leibler(KL)散度,来估计组件级量化脆弱性;随后将位分配问题表述为多选择背包问题(MCKP),从而在目标位预算下实现自适应的逐层精度分配。在ImageNet-1K数据集上针对多种视觉Transformer架构开展的大量实验表明,MixFrag在实际混合精度设置下取得了具有竞争力的分类性能。此外,在COCO目标检测与实例分割任务上的评估显示,MixFrag在现有混合精度PTQ方法中实现了最优性能,在极具挑战性的MP3/MP3设置下较此前最优方法提升了最高9.6的平均精度(AP)。额外分析验证了所提出的脆弱性度量,并证明其与学习到的位分配具有强相关性。这些结果确立了MixFrag作为视觉Transformer混合精度后训练量化的有效框架。
英文摘要
Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transformers (ViTs) on resource-constrained devices. However, existing PTQ methods typically employ uniform bit-widths across transformer components, overlooking their heterogeneous sensitivity to quantization and leading to inefficient precision allocation. In this paper, we propose {MixFrag, a fragility-guided mixed-precision PTQ framework for Vision Transformers. MixFrag first estimates component-level quantization fragility by measuring the Kullback--Leibler (KL) divergence between full-precision and isolated quantized output distributions using a small calibration set. It then formulates bit allocation as a Multiple-Choice Knapsack Problem (MCKP), enabling adaptive layer-wise precision assignment under a target bit budget. Extensive experiments on ImageNet-1K across multiple Vision Transformer architectures demonstrate that MixFrag achieves competitive classification performance under practical mixed-precision settings. Furthermore, evaluations on COCO object detection and instance segmentation show that MixFrag achieves state-of-the-art performance among existing mixed-precision PTQ methods, improving the previous best method by up to 9.6 AP under the challenging MP3/MP3 setting. Additional analyses validate the proposed fragility metric and demonstrate its strong correlation with the learned bit allocation. These results establish MixFrag as an effective framework for mixed-precision post-training quantization of Vision Transformers.