等效性错觉:大语言模型量化效应的统计表征
The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs
浏览论文内容
中文总结 AI 辅助
研究大语言模型量化效应,引入正确性一致性指标,发现适度量化下有行为差异,通过分析量化为结构算子及量化失真,揭示低比特宽度非线性断点等,促使超越传统指标评估行为。
中文摘要 AI 辅助
训练后量化广泛用于在资源受限环境中部署大语言模型,但其评估几乎完全依赖于准确性和困惑度。我们表明这些指标无法捕捉量化引起的行为变化。我们引入正确性一致性,这是一种决策级指标,用于衡量基础模型与其量化变体之间正确预测的重叠,与绝对准确性无关。在从8位到2位的多个模型和量化方案中,我们发现即使任务性能似乎保持不变,适度量化下也会出现行为差异。为了解释这种效应,我们将量化分析为注意力权重上的结构算子,并使用统计和分布度量来量化逐层失真。我们的结果揭示了低比特宽度下的非线性断点,并表明查询和键投影始终比值和输出投影更敏感。这些发现揭示了基础模型和量化模型之间的等效性错觉,并促使超越传统性能指标进行行为评估。
英文摘要
Post-Training Quantization has become widely used to compress large language models to make them deployable on resource-constrained devices. However, the evaluation of quantization methods mainly uses accuracy and perplexity, which cannot capture the behavioral changes in the quantized variants. In this work, we propose Correctness Agreement, a decision-level metric that can measure the intersection of correct predictions between the base model and its quantized variant. We use this metric across multiple models and quantization bit levels (8-bit to 2-bit), and we find that the base and quantized variants usually have a shift in behavior even when accuracy and perplexity are preserved. In order to explain this effect, we study the effect of quantization on the structure of the attention weights using statistical and distributional measures. The results reveal a breakpoint at low bit widths and show that query and key projections are more sensitive to quantization than the value and output projections. These results prove the illusion of equivalency between the base and quantized models and inspire behavioral evaluation beyond perplexity and accuracy for quantization methods.
发表机构
- University of Manitoba(曼尼托巴大学)
- Red River College Polytechnic(红河理工学院)
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。