准确性还不够:一种基于散度的方法评估量化LLM中的保真度损失
Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs
- Fraunhofer IAIS(弗劳恩霍夫智能分析与信息处理研究所)
- University of Bonn(波恩大学)
- Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所)
- University of Manitoba(曼尼托巴大学)
- University of Central Florida(中佛罗里达大学)
- Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对量化LLM评估中仅依赖准确性的不足,提出基于预测分布散度的保真度评估框架,通过JSD和TVD量化分布偏移,实验表明散度指标有效补充准确性信号。
AI中文摘要:
大型语言模型(LLM)在内存受限的边缘设备上的部署严重依赖于激进的训练后量化。然而,评估这些模型主要基于零样本任务准确性,这仅依赖于argmax预测,并且对底层预测分布的变化不敏感。因此,在逐步量化下,准确性可能表现出不稳定、非单调的行为,掩盖了相对于BFloat16(BF16)未压缩基础模型的实质性保真度损失,并提供误导性的部署信号。我们引入了一个分布敏感的评估框架,将量化LLM中的信息损失量化为标记决策边界处全词汇预测分布之间的散度。我们计算统计距离,包括Jensen-Shannon散度和总变差距离,在全精度和量化模型的输出之间,从而能够对分布偏移进行细粒度分析。使用该框架,我们量化了相对于BF16参考的概率质量位移和分布漂移,捕捉了top-1准确性未反映的预测分布变化。我们在渐进量化制度下(从未压缩的BF16到Q2_K)进行了120次运行的实验矩阵,涵盖五个基础架构和四个推理基准,提供了系统的保真度分析。我们的结果表明,在更强量化下散度指标通常增加,以保真度信号补充任务准确性。在测试的混合精度方案中,混合精度Q4_K在相似内存占用下通常比均匀Q4_0产生更低的散度。这些发现促使分布感知评估作为任务准确性的实用诊断补充;它们并不直接确立正确性、校准、安全性或用户感知质量。
英文摘要:
Deployment of Large Language Models (LLMs) on memory-constrained edge devices relies heavily on aggressive post-training quantization. However, evaluating these models is largely based on zero-shot task accuracy, which depends solely on argmax predictions and is insensitive to changes in the underlying predictive distribution. Consequently, accuracy can exhibit unstable, non-monotonic behavior under progressive quantization, masking substantial fidelity loss relative to the BFloat16 (BF16) uncompressed base model and providing misleading deployment signals. We introduce a distribution-sensitive evaluation framework quantifying information loss in quantized LLMs as the divergence between full-vocabulary predictive distributions at the token decision boundary. We compute statistical distances, including Jensen-Shannon Divergence and Total Variation Distance, between outputs of full-precision and quantized models, enabling a fine-grained analysis of distributional shift. Using this framework, we quantify probability mass displacement and distributional drift relative to the BF16 reference, capturing predictive distribution changes not reflected in top-1 accuracy. We conduct a 120-run experimental matrix across five foundation architectures and four reasoning benchmarks under progressive quantization regimes, from uncompressed BF16 to Q2_K, providing a systematic fidelity analysis. Our results show divergence metrics generally increase under stronger quantization, complementing task accuracy with a fidelity signal. Across tested llama-cpp schemes, mixed-precision Q4_K generally yields lower divergence than uniform Q4_0 at similar memory footprints. These findings motivate distribution-aware evaluation as a practical diagnostic complement to task accuracy; they do not directly establish correctness, calibration, safety, or user-perceived quality.