RAMP:面向边缘CPU视觉模型的鲁棒自适应混合精度量化
RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models
浏览论文内容
中文总结 AI 辅助
针对边缘CPU视觉模型,系统评估13种敏感性指标,提出基于Jensen-Shannon散度和K-Means聚类的鲁棒自适应混合精度量化方法,实现近乎无损精度和1.81倍加速。
中文摘要 AI 辅助
在边缘CPU上部署深度学习模型受到计算和内存限制的瓶颈制约。混合精度量化有望在保持精度的同时降低推理延迟。然而,量化对不同层类型的影响方式不一致,因此识别在何处精度损失最小且延迟降低最大至关重要,因为这种影响会在整个部署过程中累积,要么带来显著的收益,要么导致不可接受的任务性能下降。这种识别依赖于敏感性指标,即在不评估每个候选策略的任务精度的情况下,估计逐层退化程度的代理指标。然而,广泛使用的指标在现代架构上系统性地失效。我们对13种用于逐层INT8量化的敏感性指标在四个截然不同的神经网络上进行了系统的实证研究,并在两个ARM64平台上验证了由此产生的量化策略。基于梯度的敏感性方法在8个模型-硬件配置中的4个上失效,基于权重的统计方法在2个上失效。相比之下,Jensen-Shannon散度实现了零灾难性失败,可靠地隔离了无法安全量化的层。仅凭敏感性指标并不能定义策略,而通常用于该步骤的固定阈值在现代架构高度偏斜的分布上很脆弱。我们通过K-Means聚类解决了这一问题,实现了近乎无损的精度,并相对于全精度模型获得了平均1.81倍的加速。最后,我们揭示出,将加速效果可忽略不计的层(无论其敏感性如何)排除在量化之外可能适得其反,因为它会导致计算图碎片化并禁用算子融合。我们的结果为在异构边缘CPU上部署量化视觉模型的研究人员和从业者提供了具体的分配策略,而无需GPU访问或梯度计算。
英文摘要
Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. However, quantization affects different layer types in inconsistent ways, so identifying where accuracy loss is minimized and latency reduction is maximized is critical, as the effect accumulates over a full deployment into substantial savings or unacceptable task degradation. Such identification relies on sensitivity metrics, proxies that estimate layer-wise degradation without evaluating the task accuracy of every candidate policy. Nevertheless, widely used metrics fail systematically on modern architectures. We present a systematic empirical study of 13 sensitivity metrics for layer-wise INT8 quantization across four distinctly different neural networks, and validate the resulting policies on two ARM64 platforms. Gradient-based sensitivity methods fail on 4 out of 8 model-hardware configurations and weight-based statistics on 2. In contrast, the Jensen-Shannon Divergence achieves zero catastrophic failures, reliably isolating the layers that cannot be safely quantized. A sensitivity metric alone does not define a policy, and the fixed thresholds typically used for that step are fragile over the highly skewed distributions of modern architectures. We address this with K-Means clustering, achieving near-lossless accuracy and a mean speed-up of $1.81\times$ over the full-precision model. Finally, we reveal that excluding from quantization the layers whose speed-up is negligible, regardless of their sensitivity, can be counterproductive, as it induces computational graph fragmentation and disables operator fusion. Our results yield concrete allocation policies for practitioners and researchers deploying quantized vision models on heterogeneous edge CPUs, without GPU access or gradient computation.
发表机构
- Barcelona Supercomputing Center (BSC)(巴塞罗那超级计算中心)
- Universitat Politècnica de Catalunya - BarcelonaTech (UPC)(加泰罗尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。