AI 中文总结
研究大语言模型极低比特量化问题,提出跨层误差补偿和有限样本特征统计匹配联合优化方法,证明相关递归对任意非线性层精确化,实现高效实现,实验表明该方法在降低困惑度比及保持特征统计差异等方面效果显著。
AI 中文摘要
大语言模型的逐层训练后量化孤立地最小化各层的重构误差,使得量化误差在深度上累积,在极低比特情况下导致严重退化。我们将量化表述为对所有层的离散码和尺度的联合优化,由两种机制驱动:(i)跨层误差补偿,通过递归\(e_{l + 1}=A_le_l + q_l\)维持网络级累积误差,其中传播算子\(A_l\)从层的输入微分导出,局部量化残差\(q_l\)在教师特征处评估;(ii)有限样本特征统计匹配,在相对归一化下对齐全精度和量化网络之间的均值、投影协方差和中心化经验核。我们证明将传播算子实例化为量化网络的有限差分使递归对任意非线性层精确化,实现高效的前向差分实现。二元权重通过具有退火逆温度和组-wise对数尺度的镜像下降参数化\(u = \tanh(\beta z)\)进行优化。在具有1.125比特组二元权重的Qwen2.5 - 1.5B上,仅误差补偿相对于FP16教师达到9.56±0.15的困惑度比,优于逻辑蒸馏(14.09±0.53;相对32%,在3个种子上超过8个标准差)和层局部重构两个数量级。相同目标不变地转移到4比特量化(层局部为1.060对1.088)。域外评估(C4,CNN/DailyMail)表明误差补偿的优势在域外增加,而统计匹配使域外特征统计差异保持较低(0.42 - 0.88对无此情况时的1.41 - 2.99),揭示了两种机制之间的互补分工。
英文摘要
Layer-wise post-training quantization of large language models minimizes each layer's reconstruction error in isolation, allowing quantization errors to accumulate across depth and causing severe degradation in extreme low-bit regimes. We formulate quantization as a joint optimization over the discrete codes and scales of all layers, driven by two mechanisms: (i) cross-layer error compensation, which maintains the network-level accumulated error through the recursion e_{l+1} = A_l e_l + q_l, with a propagation operator A_l derived from the layer's input differential and a local quantization residual q_l evaluated at teacher features; and (ii) finite-sample feature-statistics matching, which aligns means, projected covariances, and centered empirical kernels between the full-precision and quantized networks under relative normalization. We prove that instantiating the propagation operator as a finite difference of the quantized network makes the recursion exact for arbitrary nonlinear layers, enabling an efficient forward-difference implementation. Binary weights are optimized via a mirror-descent parameterization u = tanh(beta*z) with annealed inverse temperature and group-wise log-scales. On Qwen2.5-1.5B with 1.125-bit group-binary weights, error compensation alone reaches a perplexity ratio of 9.56 +/- 0.15 over the FP16 teacher, outperforming logit distillation (14.09 +/- 0.53; 32 percent relative, more than 8 sigma over 3 seeds) and layer-local reconstruction by two orders of magnitude. The same objective transfers unchanged to 4-bit quantization (1.060 vs. 1.088 for layer-local). Out-of-domain evaluations (C4, CNN/DailyMail) show the advantage of error compensation grows off-domain, while statistics matching keeps feature-statistics discrepancy low off-domain (0.42-0.88 vs. 1.41-2.99 without it), revealing a complementary division of labor between the two mechanisms.