arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型中的量化退化:一种信号-噪声视角

Quantization Degradation in Large Language Models: A Signal-Noise Perspective

Chenxi Zhou, Pengfei Cao, Jinyu Ye, Bohan Yu, Haida Yu, Jiang Li, Jun Zhao, Kang Liu

arXiv 2608.08188首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Inner Mongolia University(中国科学院自动化研究所; 中国科学院大学; 内蒙古大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究从信号-噪声视角,系统分析了不同位宽、量化方法等因素下大语言模型量化退化的规律,揭示了退化源于源端误差产生与跨层累积的共同作用。

AI 中文摘要

后训练量化降低了大语言模型的部署成本,但量化模型的退化程度并非仅由位宽决定。我们针对多个模型族,系统研究了不同位宽、量化方法、模型规模和下游任务下的仅权值后训练量化。我们发现,这种退化程度在这些因素间存在显著差异:4位量化通常能保持性能,2位量化常导致广泛退化,而在3位时退化变得明显,但会随任务类型、量化方法和模型规模发生显著变化。为解释这种变异性,我们采用信噪比(SNR)衡量量化对全精度表示的扰动强度。我们将退化追溯到两个关联过程:单个模块内量化误差的产生方式,以及误差在各层间的累积方式。首先,源信噪比分解表明,新引入的误差取决于三个因素:权值误差的幅度、特定任务信号的强度,以及量化误差与特定任务激活的对齐程度;不同因素对这些分量的影响方式不同。其次,跨层传播分析显示,这些误差在各层传递时会被衰减、保留或放大,且更大规模的模型受益于更弱的误差放大。综上,这些结果表明,量化退化由误差在源端的产生方式及其在网络各层的累积方式共同决定。

英文摘要

Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks on multiple model families. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation, and at 3-bit, degradation becomes apparent but varies markedly with task type, quantization method and model scale. To explain this variability, we use the signal-to-noise ratio (SNR) to measure how strongly quantization perturbs full-precision representations. We trace degradation back to two linked processes: how quantization errors arise within individual modules, and how they accumulate across layers. First, a source SNR decomposition shows that newly introduced errors depend on three factors: the magnitude of the weight error, the strength of the task-specific signal, and how strongly the quantization error aligns with task-specific activations. Different factors affect these components in distinct ways. Second, a cross-layer propagation analysis shows that these errors can be attenuated, preserved, or amplified as they pass across layers, and that larger models benefit from weaker error amplification. Together, these results establish that quantization degradation is governed by how errors are introduced at the source and how they accumulate across the network.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑