发表机构
Tsinghua University; Bosch AI Research(清华大学; 博世人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探究训练后量化为何有效,发现层间误差抵消与LM头几何结构优先保留高置信词元分数是两大关键机制,解释了量化误差经多层传播仍影响甚微的原因。
AI 中文摘要
训练后量化通过以降低的精度存储大语言模型(LLM)的权重来压缩模型,每个量化权重都会在隐藏状态中引入误差。直观上,这些误差应随深度累积并破坏下一个词元的预测;随机初始化的模型会迅速累积这些差异,而量化后的预训练模型累积的隐藏状态误差要小得多,并且很大程度上保持了下游任务的性能,尽管它们从未在量化噪声下进行过训练。这引出了我们研究的问题:为什么训练后量化有效?通过比较全精度和量化前向传播,我们识别出两种表征预训练量化鲁棒性的机制。首先,某一层新引入的误差倾向于抵消该层从其输入继承的误差。两者部分抵消,使得全精度与量化传播之间的差异增长缓慢。这种对抗性残差交互在预训练期间形成。我们的定量分析将其识别为减缓隐藏误差增长的主要因素。其次,LM头(语言模型头)的几何结构优先保留排名靠前词元的分数和概率,这些词元通常代表模型最自信的预测。这两种机制共同解释了为什么经过多层传播的量化误差仍只能产生较小的输出变化,并且我们在多种模型和量化设置下验证了这些发现。
英文摘要
Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and corrupt next-token prediction; randomly initialized models accumulate these discrepancies rapidly, whereas quantized pretrained models accumulate much less hidden-state error and largely maintain downstream task performance, even though they were never trained with quantization noise. This raises the question we address: why does post-training quantization work? Comparing full-precision and quantized forward passes, we identify two mechanisms that characterize pretrained quantization robustness. First, the error a layer newly introduces tends to oppose the error it inherits from the layer's input. The two cancel partially such that the discrepancy between full-precision and quantized passes grows slowly. This counteracting residual interaction develops during pretraining. Our quantitative analysis identifies it as a major factor slowing hidden-error growth. Second, LM-head geometry preferentially preserves the scores and probabilities of high-ranked tokens, which typically represent the model's most confident predictions. Together, these mechanisms explain why quantization error that passes through numerous layers can still produce only small output changes, and we verify the findings across models and quantization settings.
Comments45 pages, 26 figures, including appendices