发表机构
Intel Labs China(英特尔中国研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示LLM量化中MSE重建损失导致跨阶段优化失衡,提出优化强度应与损失尺度解耦,并证明RMSE变体通过隐式梯度归一化显著提升PTQ性能。
AI 中文摘要
训练后量化(PTQ)方法通常采用顺序量化,将预训练的大语言模型(LLM)划分为一系列单元(如Transformer块),每个阶段量化一个单元。最先进的PTQ方法主要基于学习,通过梯度下降优化辅助量化参数(如缩放因子、旋转矩阵、裁剪阈值和适配器),以最小化重建损失。常见做法是使用均方误差(MSE)作为重建损失函数,但其引发的优化行为在很大程度上尚未被探索。在本工作中,我们对顺序量化采取整体视角,系统研究优化从第一个量化阶段到最后一个阶段如何演变,旨在深入理解基于学习的PTQ方案中的优化过程。通过涵盖代表性基于学习的PTQ方法、LLM系列、模型规模、架构、量化设置和各种任务的广泛实证研究,我们一致发现优化失衡:重建损失幅度在各阶段间差异巨大,伴随MSE下高度不均匀的梯度幅度和参数更新。我们将跨阶段的损失幅度范围称为重建损失尺度,并揭示MSE将异常大的重建损失尺度转化为高度不均匀的梯度幅度,进而导致各量化阶段优化强度不均。这一发现提出了改进基于学习的PTQ的一般原则:各阶段的优化强度应与重建损失尺度解耦。理论上,我们证明在样本、通道、标记和元素级别定义的均方根误差(RMSE)变体通过隐式梯度归一化自然实现这一原则,作为直接替换显著优于MSE。
英文摘要
Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ methods are predominantly learning-based, optimizing auxiliary quantization parameters (e.g., scaling factors, rotation matrices, clipping thresholds, and adapters) via gradient descent to minimize a reconstruction loss. A common practice is to use mean squared error (MSE) as the reconstruction loss function, yet its induced optimization behavior remains largely unexplored. In this work, we take a holistic view of sequential quantization and systematically investigate how optimization evolves from the first quantization stage to the last, aiming for a deep understanding of optimization in learning-based PTQ schemes. Through extensive empirical studies spanning representative learning-based PTQ methods, LLM families, model scales, architectures, quantization settings and various tasks, we consistently uncover Optimization Imbalance: reconstruction loss magnitudes vary dramatically across stages, accompanied by highly uneven gradient magnitudes and parameter updates under MSE. We term the cross-stage range of loss magnitudes the reconstruction loss scale, and reveal that MSE translates the unexpectedly large reconstruction loss scale into highly uneven gradient magnitudes, which in turn lead to uneven optimization strength across quantization stages. This finding suggests a general principle for improving learning-based PTQ: optimization strength across stages should be decoupled from the reconstruction loss scale. Theoretically, we show that root mean squared error (RMSE) variants defined at the sample, channel, token, and element levels naturally realize this principle through implicit gradient normalization, outperforming MSE significantly as a drop-in replacement.
CommentsProject page: https://github.com/IntelChina-AI/RMSE