低维高杠杆子空间优化:超越神经网络量化的全参数耦合训练
Low-Dimensional High-Leverage Subspace Optimization: Beyond Full-Parameter Coupled Training for Neural Network Quantization
AI总结:
本研究针对神经网络量化中全参数耦合训练的缺陷,提出NAP方法,通过针对性优化归一化仿射参数子空间,在PTQ和QAT场景下提升低比特量化性能,在ImageNet等数据集上验证了其有效性。
AI中文摘要:
低比特量化在紧凑网络上会遭受严重的精度下降,这源于主导性的全参数耦合训练范式,该范式忽略了参数子空间的异质性,紧凑网络的特征冗余有限,几乎没有空间吸收量化误差。传统流程采用整体优化:PTQ(后训练量化)重构固定的预训练模型,不提升固有量化友好性;QAT(量化感知训练)联合更新所有参数,会遭受骨干权重与校准参数之间的梯度耦合。在本文中,我们识别出归一化仿射参数为主导量化鲁棒性的低维高杠杆子空间,并提出归一化仿射预处理(Normalization Affine Preconditioning,NAP)用于针对性子空间优化。对于PTQ,NAP冻结骨干权重,仅在全精度模型的目标伪量化图上微调仿射参数,在下游重构前主动提升量化友好性;对于QAT,我们引入交替QAT-NAP方案,将特征学习与数值校准解耦,突破饱和联合训练的性能上限。理论分析证实,BN(批归一化)仿射参数可完全抵消量化失真的逐通道仿射分量,而非线性舍入与截断残差构成不可约误差边界;蒸馏引导的NAP作为定向平坦度优化,将师生logit(对数几率)失配投影到受限子空间。在ImageNet和CIFAR-100上的实验表明,NAP可恢复严重崩溃的低比特量化,持续提升基于重构的PTQ,且以可忽略的调优成本超越饱和的全参数QAT。本研究揭示了针对性低维子空间优化的原理,为高效深度学习提供了超越全参数耦合训练的新视角。
英文摘要:
Low-bit quantization suffers severe accuracy degradation on compact networks, rooted in the dominant full-parameter coupled training paradigm that ignores parameter subspace heterogeneity. Their limited feature redundancy leaves little room to absorb quantization errors. Conventional pipelines adopt monolithic optimization: PTQ reconstructs fixed pretrained models without improving inherent quantization friendliness; QAT updates all parameters jointly, suffering from gradient coupling between backbone weights and calibration parameters. In this paper, we identify normalization affine parameters as a low-dimensional high-leverage subspace dominating quantization robustness, and propose Normalization Affine Preconditioning (NAP) for targeted subspace optimization. For PTQ, NAP freezes backbone weights and fine-tunes only affine parameters under the target fake-quantization graph on full-precision models, proactively boosting quantization friendliness before downstream reconstruction. For QAT, we introduce an alternating QAT-NAP schema that decouples feature learning and numerical calibration, breaking the performance ceiling of saturated joint training. Theoretical analysis confirms BN affine parameters fully cancel the channel-wise affine component of quantization distortion, while nonlinear rounding and clipping residuals form the irreducible error boundary; distillation-guided NAP acts as directional flatness optimization, projecting teacher-student logit mismatch onto the restricted subspace. Experiments on ImageNet and CIFAR-100 show NAP recovers severely collapsed low-bit quantization, consistently boosts reconstruction-based PTQ, and outperforms saturated full-parameter QAT with negligible tuning cost. This work reveals the principle of targeted low-dimensional subspace optimization, offering a new perspective beyond full-parameter coupled training for efficient deep learning.