arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21834math.OC

超越影子权重:将量化感知训练视为量化端点下降

Beyond Shadow Weights: Quantization-Aware Training as Quantized-Endpoint Descent

  • University of California, Berkeley(加州大学伯克利分校)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Sheng-An Xu, Hanyang Li, Jianhao Ma, Ying Cui

AI总结:

本文提出将量化感知训练视为量化端点下降,并基于理想化模型提出QAR算法,直接优化量化码,在降低内存成本的同时达到与QAT相当或更优的性能。

AI中文摘要:

量化感知训练(QAT)更新全精度影子权重 $\mathbf{x}$,但部署量化端点 $Q(\mathbf{x})$。现有对QAT的解释主要从影子权重的角度出发:QAT可以将 $\mathbf{x}$ 移向更平坦的盆地,从量化引起的振荡中获得鲁棒性,或在影子损失 $f(\mathbf{x})$ 与量化误差 $\\|\mathbf{x}-Q(\mathbf{x})\\|_2$ 之间取得平衡。这些视角并不能直接解释一个经验观察:部署端点损失 $f(Q(\mathbf{x}))$ 得到改善,而影子损失 $f(\mathbf{x})$ 并未改善,甚至可能大幅增加。在本文中,我们通过将QAT视为有限网格端点动力学来提供一种不同的解释。受缩放预训练权重近似正态性的启发,我们为残差阶段提出了一个理想化模型,该模型记录每个影子权重在其量化单元内的位置,作为单元宽度的分数。该模型导出一个交叉定律,用于确定在影子更新后哪些坐标会跨越量化边界。受理想化模型和QAT中信号不平衡现象的启发,我们进一步提出了QAR(带放大路由的量化),一种直接作用于量化码的算法框架。与QAT相比,QAR既有理论依据又节省内存:对于一族功率放大器,它允许在不可避免的有限网格下限内获得可行梯度界,而无需保留全精度影子权重副本。在大语言模型后训练上的实验提供了与端点观点一致的证据,并表明QAR在更小的内存成本下可以达到与QAT相当或更好的性能。

英文摘要:

Quantization-aware training (QAT) updates a full-precision shadow weight $\mathbf{x}$ but deploys the quantized endpoint $Q(\mathbf{x})$. Existing explanations for QAT largely view its success through the lens of shadow weights: QAT can move $\mathbf{x}$ toward flatter basins, gain robustness from quantization-induced oscillations, or balance the shadow loss $f(\mathbf{x})$ against the quantization error $\|\mathbf{x}-Q(\mathbf{x})\|_2$. These perspectives do not directly explain the empirical observation that the deployed endpoint loss $f(Q(\mathbf{x}))$ improves while the shadow loss $f(\mathbf{x})$ does not, and can even increase substantially. In this paper, we offer a different explanation by treating QAT as finite-grid endpoint dynamics. Motivated by the approximate normality of rescaled pretrained weights, we propose an idealized model for the residual phase, which records where each shadow weight sits inside its quantization cell as a fraction of the cell width. This model leads to a crossing law that determines which coordinates cross quantization boundaries after a shadow update. Inspired by the idealized model and signal-imbalance phenomenon in QAT, we further propose QAR (Quantization with Amplified Routing), an algorithmic framework that directly operates on the quantization code. In contrast to QAT, QAR is both theoretically grounded and memory-efficient: it admits feasible-gradient bounds for a family of power amplifiers up to unavoidable finite-grid floors without retaining a full-precision shadow weight copy. Experiments on post-training of large language models provide evidence consistent with the endpoint view and show that QAR can be comparable to or better than QAT with smaller memory cost.

↑