发表机构
ZTE Corporation; The Chinese University of Hong Kong; Northwestern Polytechnical University; Peking University; Nanjing University of Aeronautics and Astronautics(中兴通讯; 香港中文大学; 西北工业大学; 北京大学; 南京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
G$^2$PTQ提出一种统一的后训练量化框架,通过广义梯度补偿整合一阶和二阶信息,并采用信任区域缩放机制稳定梯度,在多种模型和位宽上优于现有基线。
AI 中文摘要
后训练量化(PTQ)是一种无需重新训练即可减少大型语言模型(LLM)内存和计算占用量的实用方法。基于GPTQ的方法已成为事实上的标准,但它们存在两个互补的局限性。具有局部、逐层目标的方法缺乏全局监督;而具有全局目标的方法在开始时固定其Hessian估计并忽略一阶梯度,因此随着量化的进行,其指导会变得过时。本文提出了G$^2$PTQ,一个统一的PTQ框架,采用广义梯度补偿,在全局监督的块级优化目标下整合一阶和二阶信息。通过在量化每个Transformer块之前刷新梯度和Hessian估计,G$^2$PTQ避免了先前全局方法的过时问题。此外,为了稳定精确的一阶补偿,我们引入了一种信任区域缩放机制,动态限制梯度步长以防止权重更新爆炸。最后,我们推导了块级Hessian近似和精确梯度补偿的高效实现。在多种模型家族和位宽上的实验结果表明,G$^2$PTQ能够更好地与全精度模型对齐,优于最先进的基线方法。代码可在以下网址获取:此https URL。
英文摘要
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds. This paper presents G$^2$PTQ, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective. By refreshing gradient and Hessian estimates before quantizing each Transformer block, G$^2$PTQ avoids the staleness of prior global methods. Furthermore, to stabilize the exact first-order compensation, we introduce a trust-region scaling mechanism that dynamically bounds the gradient step to prevent exploding weight updates. Finally, we derive efficient implementations for block-wise Hessian approximation and exact gradient compensation. Experimental results on various model families and bit-widths demonstrate that G$^2$PTQ enables better alignment with the full-precision model, outperforming state-of-the-art baselines. Code is available at: https://github.com/G2PTQ/G2PTQ.