arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于随机梯度下降优化模型的低位后训练量化前的高效调优

Efficient Tuning Before Low-Bit Post-Training Quantization for Stochastic Gradient Descent-optimized Models

Peng Xia, Junbiao Pang, Muhammad Ayub Sabir

arXiv 2607.11359首次发表:更新:

发表机构

School of Information Science and Technology, Beijing University of Technology(北京工业大学信息科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对低位后训练量化性能易降问题,提出ETBQ方法,在随机梯度下降优化模型的PTQ前进行预处理调优,通过从量化误差分布采样扰动优化全精度模型,实验证明该方法能提升不同任务的低位PTQ性能。

AI 中文摘要

后训练量化(PTQ)用于在有限内存和计算预算下压缩深度神经网络以进行部署。然而,低位(如2位或4位)PTQ往往会导致性能大幅下降。大多数现有PTQ方法在无约束的全精度(FP)模型上运行,主要通过事后重建解决量化误差。本文提出量化前的高效调优(ETBQ),这是一个在PTQ之前用于随机梯度下降(SGD)优化模型的预处理调优阶段。在调优期间,FP模型在从权重和激活量化的误差分布中采样的扰动下进行优化,引导模型走向对后续PTQ不太敏感的损失景观区域。实验表明ETBQ在不同任务中均能提高低位PTQ的性能。

英文摘要

Post-training quantization (PTQ) compresses deep neural networks for deployment under limited memory and computational budgets. However, low-bit (i.e., 2-bit or 4-bit) PTQ often suffers from substantial performance degradation. Most existing PTQ methods operate on an unconstrained full-precision (FP) model and primarily address quantization errors through post-hoc reconstruction. We argue that low-bit PTQ accuracy is limited not only by post-quantization error minimization, but also by the quantization-error tolerance of a FP model itself. In this paper, we propose Efficient Tuning Before Quantization (ETBQ), a pre-conditioning tuning stage for Stochastic Gradient Descent (SGD)-optimized models before PTQ. During tuning, the FP model is optimized under perturbations sampled from the error distributions of weight and activation quantization, guiding the model toward a loss-landscape region that is less sensitive to the subsequent PTQ. Unlike QAT, ETBQ does not train a fake-quantized deployment model, which is computationally and memory intensive. Instead, ETBQ outputs a FP model that can be used by any PTQ backend. Experiments on CIFAR-100, Tiny-ImageNet, ImageNet, and Cityscapes provide consistent evidence that ETBQ improves low-bit PTQ across diverse tasks. Under W2A4 settings, e.g., ETBQ improves over naive PTQ by 2.14\% top-1 accuracy on Tiny-ImageNet and by 5.80\% mIoU on Cityscapes. Code is available at https://github.com/xpxpxp2001xpxpxp/ETBQ.

Commentsv2 revision: Added hyperparameter settings of all experiments in appendix, fixed minor typos, adjusted figure layout, polished experimental analysis. 12 pages, 10 figures, submitted to IEEE Transactions on Neural Networks and Learning Systems (TNNLS). Code available at https://github.com/xpxpxp2001xpxpxp/ETBQ

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑