AI 中文总结
研究大语言模型PTQ在低精度下降问题,提出RDQ框架,通过级联误差补偿,在多种测试架构上取得最优结果,输出标准量化且零运行时开销。
AI 中文摘要
大语言模型的训练后量化(PTQ)在低于4位精度时性能急剧下降。我们发现根本原因是残差流分布漂移:每个Transformer层注入的量化噪声在共享残差表示中累积,导致与FP16基线的KL散度随深度超线性增长。我们发现84%的LLaMA-3-8B层呈现非高斯残差分布,且每层残差流方差在深度上增长6548倍。我们提出RDQ框架,其核心贡献是级联误差补偿(CEC)。在三种测试架构上,RDQ取得了最优结果,输出为标准的组128非对称量化,零运行时开销。
英文摘要
Post-training quantization (PTQ) of large language models degrades sharply below 4-bit precision. We identify the root cause as residual stream distributional drift: quantization noise injected at each transformer layer accumulates in the shared residual representation, causing KL divergence from the FP16 baseline to grow super-linearly with depth (Pearson r=0.999 with log-perplexity, p<0.001, confirmed across all tested methods and bit-widths). We discover that 84% of LLaMA-3-8B layers exhibit non-Gaussian residual distributions (KS test, p<=0.05), and that per-layer residual stream variance grows 6,548x across depth. We propose RDQ (Residual Distribution Quantization), a PTQ framework whose central contribution is Cascaded Error Compensation (CEC): a sequential calibration procedure that captures the actual drifted activations each layer receives (computed by running calibration data through already-quantized upstream layers) and fits per-channel AWQ-style scales against those drifted inputs, with scales folded into preceding RMSNorm weights for exact mathematical equivalence at zero inference overhead. RDQ achieves state-of-the-art results on all three tested architectures: LLaMA-3-8B: 7.55 / 5.62 PPL (W3/W4); Qwen-2.5-7B: 7.46 / 6.38 PPL; Mistral-7B: 6.88 / 5.73 PPL. RDQ beats the best published baseline (LeanQuant/SpinQuant) at every model and bit-width combination, with gains up to -46.4% vs. RTN at W3A16 on LLaMA-3-8B. All output is standard group-128 asymmetric quantization, deployable on Qualcomm AIMET, GGUF, and any standard inference stack at zero runtime overhead.
CommentsI am withdrawing due to compliance issue of my organisation