SCULPT:训练边缘视觉模型以实现训练后量化就绪
SCULPT: Training Edge Vision Models for Post-Training Quantization Readiness
浏览论文内容
中文总结 AI 辅助
提出SCULPT方法,在FP32微调时抑制激活分布的不利统计特性,学习可直接用于PTQ的裁剪边界,无需QAT或事后修复,实现边缘视觉模型的低位宽部署。
中文摘要 AI 辅助
边缘视觉模型难以部署在资源受限的硬件上,这使得低位宽训练后量化(PTQ)颇具吸引力。实际中,标准FP32训练通常会产生重尾激活分布,其离群值会破坏激活量化:保留全范围会在罕见极值上浪费量化区间,而激进裁剪则会导致信息丢失。现有解决方案通常依赖量化感知训练(QAT),这会增加训练复杂度并存在位宽耦合问题,或依赖训练后修复模型的高级PTQ流程。我们提出SCULPT(Statistical Clipping and Uniform Loss for Post-Training,即用于训练后的统计裁剪与均匀损失),这是一种在常规FP32微调期间提升PTQ就绪性的训练时方法。SCULPT结合了拓扑感知激活正则化器(用于抑制对量化不利的偏度和峰度)与稳定的基于百分位数的裁剪机制(用于学习可部署的激活边界)。与QAT不同,SCULPT在优化过程中不模拟量化;与事后离群值修复的PTQ方法不同,它不需要运行时激活变换。学习到的裁剪边界可直接导出到标准PTQ流程中,用于低位宽部署,包括INT8及如W4A8这类低位宽设置。
英文摘要
Edge vision models are difficult to deploy on resource-constrained hardware, making low-bit post-training quantization (PTQ) attractive. In practice, standard FP32 training often produces heavy-tailed activation distributions whose outliers destabilize activation quantization: preserving the full range wastes quantization bins on rare extremes, while aggressive clipping causes information loss. Existing solutions typically rely on quantization-aware training (QAT), which adds training complexity and bit-width coupling, or advanced PTQ procedures that repair the model after training. We present SCULPT (Statistical Clipping and Uniform Loss for Post-Training), a training-time method that improves PTQ readiness during ordinary FP32 fine-tuning. SCULPT combines a topology-aware activation regularizer that suppresses quantization-hostile skewness and kurtosis with a stable percentile-based clipping mechanism that learns deployment-ready activation bounds. Unlike QAT, SCULPT does not simulate quantization during optimization; unlike post hoc outlier-repair PTQ methods, it does not require runtime activation transformations. The learned clipping bounds can be exported directly into a standard PTQ workflow for low-bit deployment, including INT8 and lower-bit settings such as W4A8.
发表机构
- Valeo Vision Systems(法雷奥视觉系统)
- Valeo India Pvt. Ltd(法雷奥印度私人有限公司)
机构由 AI 辅助整理,请以论文原文为准。