发表机构
Technical University of Munich; Microsoft Research Cambridge(慕尼黑工业大学; 微软研究院剑桥分院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出激活去噪的并行量化方法,通过鲁棒性正则化模拟串行量化优势,在保持并行高效的同时减少误差累积,显著提升LLM量化精度与速度。
AI 中文摘要
训练后量化是压缩大型语言模型的有力工具。最具可扩展性的方法并行量化每一层,但量化误差随后会通过残差流累积,因为没有层会纠正其前面层的误差。串行量化通过在其前驱的已量化输出上重新校准每一层来考虑这种误差累积,从而产生更强的结果,但代价是串行调度在大规模下成为瓶颈。作为解决方案,我们提出了带有激活去噪的并行量化,它在保持量化完全并行的同时,恢复了许多串行量化的好处。我们不是逐层重新校准,而是采取鲁棒性视角,将上游误差建模为噪声,通过预处理步骤和度量加权舍入进行正则化以对其鲁棒。应用于每一层时,这种正则化形成了一种深度累积的平滑性惩罚,抑制了量化误差通过模型的放大。与量化中常用的正交旋转(必须保持模型功能)不同,我们将权重乘以更一般的线性变换。我们发现这两者是互补的,其效果会累积。实验上,我们的鲁棒性正则化在单次并行过程中恢复了串行量化收益的很大一部分,而时间仅为其一小部分。总体而言,通过将累积量化误差视为鲁棒性问题,我们为大规模更高效、更准确的LLM量化提供了原则性基础。
英文摘要
Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but quantization errors then compound through the residual stream, as no layer corrects for the errors of the layers before it. Sequential quantization accounts for this error compounding by re-calibrating each layer on the already-quantized outputs of its predecessors, yielding stronger results but at the cost of a serial schedule that becomes a bottleneck at scale. As a solution, we propose parallel quantization with activation denoising, which recovers much of the sequential benefit while keeping quantization fully parallel. Rather than re-calibrating layer-by-layer, we take a robustness perspective and model the upstream error as noise, regularizing to be robust to it through a preprocessing step followed by metric-weighted rounding. Applied at every layer, this regularization forms a depth-compounding smoothness penalty that dampens how strongly quantization errors amplify through the model. Unlike orthogonal rotations commonly used in quantization, which must preserve the model's function, we multiply the weights by a more general linear transformation. We find that the two are complementary and their effects compound. Empirically, our robustness regularization recovers a significant part of sequential quantization's benefit in a single parallel pass, at a fraction of its time. Overall, by treating compounding quantization errors as a robustness problem, we offer a principled foundation for more efficient and accurate LLM quantization at scale.