通过激活引导补偿和正交残差理解LLM量化
Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals
浏览论文内容
中文总结 AI 辅助
本文通过将量化误差分解为激活引导补偿和正交残差,推导出旋转、符号选择和缩放的实际指南,在八个Llama和Mistral模型上实现与SpinQuant竞争的性能。
中文摘要 AI 辅助
训练后权重激活量化降低了大型语言模型的内存和推理成本,但激进的W4A4量化仍然困难,因为激活异常值降低了有效量化分辨率。尽管权重优化、通道级缩放和正交旋转缓解了这一问题,但它们所解决的误差分量及其关系仍不清楚。通过将局部权重激活量化误差精确分解为激活引导的权重补偿项和正交残差,我们利用持久通道级异常值和常规激活量来界定残差。这种分解阐明了哪些误差分量可以通过权重补偿解决,哪些需要变换设计。然后,我们利用残差界推导出应用随机Hadamard旋转、符号选择和通道缩放的实用指南。特别是,分析解释了随机符号如何抑制持久异常值通道间的相长干涉,采样多个符号模式如何改善变换选择,以及二阶矩平衡如何导致$L_2$缩放规则,而进一步放宽则恢复SmoothQuant风格的$L_\infty$缩放。我们通过无反向传播配置在八个Llama和Mistral模型上评估这些指南,获得了与梯度训练的SpinQuant相竞争的性能。
英文摘要
Post-training weight-activation quantization reduces the memory and inference costs of large language models, but aggressive W4A4 quantization remains difficult because activation outliers degrade effective quantization resolution. Although weight optimization, channel-wise scaling, and orthogonal rotation mitigate this problem, the error components they address and their relationship remain unclear. Using an exact decomposition of local weight-activation quantization error into an activation-guided weight compensation term and an orthogonal residual, we bound the residual using persistent channel-wise outlier and regular activation quantities. This decomposition clarifies which error components can be addressed by weight compensation and which require transformation design. We then use the residual bounds to derive practical guidelines for applying randomized Hadamard rotation, sign selection, and channel scaling. In particular, the analysis explains how random signs suppress constructive interference among persistent outlier channels, how sampling multiple sign patterns can improve transformation selection, and how second-moment balancing leads to an $L_2$ scaling rule while a further relaxation recovers SmoothQuant-style $L_\infty$ scaling. We evaluate these guidelines through backpropagation-free configurations across eight Llama and Mistral models, obtaining performance competitive with gradient-trained SpinQuant.
发表机构
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。