AHCQ-SAM:迈向准确且兼容硬件的后训练分割任何模型量化
AHCQ-SAM: Toward Accurate and Hardware-Compatible Post-Training Segment Anything Model Quantization
- Department of Electronics and Electrical Engineering, Keio University(庆应义塾大学电子与电气工程系)
- School of Computer Science and Technology, Hainan University(海南大学计算机科学与技术学院)
- Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出AHCQ-SAM框架,通过四种协同组件解决SAM量化中的关键挑战,实现在4位SAM-B和SAM2-Tiny上的mAP和J&F指标提升,并在FPGA上验证了其硬件兼容性和能效优势。
AI中文摘要:
片段任何模型(SAM)通过其强大的零样本能力革新了图像和视频分割。然而,其庞大的参数规模和高计算需求阻碍了在资源受限的边缘设备上的高效部署。虽然后训练量化(PTQ)提供了一个实用的解决方案,但现有方法仍无法处理四个关键的量化挑战:(1)病态权重;(2)偏斜且长尾的后GELU激活;(3)线性投影中的显著通道方差;(4)指数级扩展和异质的注意力分数。为缓解这些瓶颈,我们提出了AHCQ-SAM,一个准确且硬件兼容的PTQ框架,包含四个协同组件:(1)激活感知条件数减少(ACNR),通过近端点算法正则化权重矩阵以抑制病态;(2)混合对数均匀量化(HLUQ),结合二的幂和均匀量化器以捕捉偏斜的后GELU激活;(3)通道感知分组(CAG),将具有相似统计特征的通道分组以在最小硬件开销下实现高精度;(4)对数非线性量化(LNQ),利用对数变换以自适应调整量化分辨率以适应指数和异质的注意力分数。实验结果表明,AHCQ-SAM在SAM上优于现有方法。与最先进方法相比,它在COCO数据集上4位SAM-B上实现了15.2%的mAP提升。此外,我们建立了SAM2的PTQ基准,其中AHCQ-SAM在SA-V测试数据集上4位SAM2-Tiny上实现了14.01%的J&F提升。最后,基于FPGA的实现验证了AHCQ-SAM的实用性,相比浮点基线,实现了7.12倍的速度提升和6.62倍的能效提升。
英文摘要:
The Segment Anything Model (SAM) has revolutionized image and video segmentation with its powerful zero-shot capabilities. However, its massive parameter scale and high computational demands hinder efficient deployment on resource-constrained edge devices. While Post-Training Quantization (PTQ) offers a practical solution, existing methods still fail to handle four critical quantization challenges: (1) ill-conditioned weights; (2) skewed and long-tailed post-GELU activations; (3) pronounced inter-channel variance in linear projections; and (4) exponentially scaled and heterogeneous attention scores. To mitigate these bottlenecks, we propose AHCQ-SAM, an accurate and hardware-compatible PTQ framework featuring four synergistic components: (1) Activation-aware Condition Number Reduction (ACNR), which regularizes weight matrices via a proximal point algorithm to suppress ill-conditioning; (2) Hybrid Log-Uniform Quantization (HLUQ), which combines power-of-two and uniform quantizers to capture skewed post-GELU activations; (3) Channel-Aware Grouping (CAG), which clusters channels with homogeneous statistics to achieve high accuracy with minimal hardware overhead; and (4) Logarithmic Nonlinear Quantization (LNQ), which utilizes logarithmic transformations to adaptively adjust quantization resolution for exponential and heterogeneous attention scores. Experimental results demonstrate that AHCQ-SAM outperforms current methods on SAM. Compared with the SOTA method, it achieves a 15.2% improvement in mAP for 4-bit SAM-B with Faster R-CNN on the COCO dataset. Furthermore, we establish a PTQ benchmark for SAM2, where AHCQ-SAM yields a 14.01% improvement in J&F for 4-bit SAM2-Tiny on the SA-V Test dataset. Finally, FPGA-based implementation validates the practical utility of AHCQ-SAM, delivering a 7.12x speedup and a 6.62x power efficiency improvement over the floating-point baseline.