量化感知预训练与约束经验权重分布
Quantization-Aware Pre-Training with Constrained Empirical Weight Distribution
浏览论文内容
中文总结 AI 辅助
提出无超参数、无内存开销的CEWT方法,通过将权重投影到零均值高斯分布约束点来抑制量化感知预训练中的舍入边界振荡,在增加4%训练时间下平均降低困惑度2.5点,最高21点,适用于低至1位精度的LLaMA/GPT模型。
中文摘要 AI 辅助
量化感知预训练(QAPT)可以提高深度神经网络的推理效率,但一种称为舍入边界权重振荡的问题行为会在训练过程中引入有害噪声,并显著降低收敛速度。虽然现有方法可以减少这种有害噪声,但它们要么引入额外的超参数或内存开销,要么无法持续提高模型精度。在这项工作中,我们提出了约束经验权重分布(CEWT)优化方法,这是第一种无超参数、无内存开销的振荡抑制方法,能够持续改进QAPT性能:优化器更新后步骤将权重投影到权重空间中最近的点,该点的经验分布(直方图)匹配零均值高斯分布。我们的关键洞察是,许多量化器在设计时隐含假设待量化数据是来自零均值高斯分布的样本的排列,而这一假设在QAPT期间并不成立。通过将零均值高斯先验作为硬约束强制执行,CEWT可以抑制这种有害噪声。在多种最先进量化器和超球优化器的组合上的实证结果表明,在训练时间几何平均增加4%的情况下,CEWT可以持续降低低精度(低至1位激活和权重,高达6.1亿参数)LLaMA/GPT模型的预训练困惑度(平均降低2.5点,最高降低21点),且不引入任何超参数或存储开销。代码可在以下网址获取:此https URL
英文摘要
Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this detrimental noise, they either introduce additional hyperparameters or memory overhead, or cannot consistently improve model accuracy. In this work, we propose optimization with $\textbf{C}$onstrained $\textbf{E}$mpirical $\textbf{W}$eight dis$\textbf{T}$ribution (CEWT), the first hyperparameter-free memory-overhead-free oscillation suppression method that consistently improves QAPT performance: an optimizer post-update step that projects weights to the nearest point in weight space whose empirical distribution (histogram) matches a zero-mean Gaussian. Our key insight is many quantizers are designed with the implicit assumption that the to-be-quantized data are permutations of samples from a zero-mean Gaussian, and this assumption is not true during QAPT. By enforcing the zero-mean Gaussian prior as a hard constraint, CEWT can suppress this detrimental noise. Empirical results on various combinations of SOTA quantizers and hypersphere optimizers suggest, that with a geomean increase of 4% in training time, CEWT can consistently reduce the pre-training perplexity (by an average of 2.5 and up to 21 points) of low-precision (down to 1-bit activations and weights and up to 610M parameters) LLaMA/GPT models without introducing any hyperparameters or storage overhead. Code is available at https://github.com/1733116199/cewt