arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对齐,然后修正:面向极度量化大语言模型的无训练两阶段低秩补偿

Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models

Seobin Song, Geonho Lee, Janghwan Lee, Jungwook Choi

arXiv 2610.08164首次发表:更新:

发表机构

Hanyang University(汉阳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对极度量化大语言模型,提出无训练两阶段低秩补偿框架:先非对称对齐压缩目标,再自然梯度吸收一阶残差,显著降低困惑度并提升零样本性能。

AI 中文摘要

低秩量化误差补偿(LQEC)通过在冻结的量化权重旁附加一个闭式秩-$r$适配器,无需任何训练即可恢复在激进权重量化下损失的精度。我们表明,现有补偿器受限于两个共同的简化假设。它们对称地进行校准,即在相同的激活上评估全精度权重和补偿后的权重,这导致补偿目标本质上是高秩的——因此固定的秩预算只能捕获其中的一小部分。此外,它们仅最小化损失的二阶项,尽管补偿后的模型并非处于平稳状态:每一层中仍存在一个大于所施加补偿本身的一阶下降方向,且没有任何重构目标能够吸收它。我们提出一个两阶段闭式框架,消除了这两个简化。第一阶段在Fisher加权的非对称目标下,将每层的输出与全精度模型对齐,将秩预算集中在可压缩秩的目标上。第二阶段在补偿后的模型上重新测量统计量,并应用一个秩约束的自然梯度步骤,以吸收剩余的一阶信号。每个适配器都是单次截断SVD的结果;反向传播仅用于收集统计量。在QuIP#下的2比特设置中,我们的方法将Qwen3-8B在WikiText-2上的困惑度从12.43降至10.26,将Qwen3-4B从21.11降至13.22。在留出的C4语料库上,它恢复了与FP16差距的51%和84%,而最强基线分别为31%和63%,在七任务零样本平均值、更高比特宽度以及不同量化器下均获得一致提升。

英文摘要

Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training. We show that existing compensators are limited by two shared simplifications. They calibrate symmetrically, evaluating the full-precision and compensated weights on the same activation, which yields a compensation target that is inherently high-rank -- so a fixed rank budget captures only a small fraction of it. And they minimize only the second-order term of the loss, although the compensated model is not stationary: a first-order descent direction larger than the applied compensation itself remains in every layer, and no reconstruction objective can absorb it. We propose a two-stage closed-form framework that removes both simplifications. Stage 1 aligns each layer's output with the full-precision model under a Fisher-weighted asymmetric objective, concentrating the rank budget on a rank-compressible target. Stage 2 re-measures statistics on the compensated model and applies a rank-constrained natural-gradient step that absorbs the remaining first-order signal. Every adapter is the result of a single truncated SVD; backward passes serve only to collect statistics. At 2 bits under QuIP#, our method reduces WikiText-2 perplexity from 12.43 to 10.26 on Qwen3-8B and from 21.11 to 13.22 on Qwen3-4B. On the held-out C4 corpus, it recovers 51% and 84% of the gap to FP16, versus 31% and 63% for the strongest baseline, with consistent gains in the seven-task zero-shot average, at higher bit-widths, and under a distinct quantizer.

Comments17 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑