arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化大语言模型的可靠性缩放定律

Reliability Scaling Laws for Quantized Large Language Models

Sirine Ayadi, Sándor Daróczi, Stephan Günnemann, Bertrand Charpentier

arXiv 2607.10855首次发表:更新:

发表机构

Technical University of Munich; Munich Data Science Institute; Pruna AI(慕尼黑工业大学; 慕尼黑数据科学研究所; Pruna人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究量化大语言模型的可靠性缩放定律,通过对其不确定性、校准和鲁棒性进行全面评估,刻画可靠性与模型比特总数的关系,发现4比特量化模型有可靠性峰值,且量化增强了模型对自然输入扰动的鲁棒性。

AI 中文摘要

量化是通过减少参数的比特宽度来构建高效且资源节约型大语言模型(LLM)的强大策略。虽然量化的LLM在使用标准预测指标的无扰动输入上取得了先进性能,但在使用可靠性指标衡量的扰动输入上的性能仍未得到充分探索。为填补这一空白,我们首先对量化的LLM进行了全面的可靠性评估,包括不确定性、校准和鲁棒性三个关键部分。然后,我们刻画了可靠性如何随模型比特总数缩放。研究表明,性能随比特总数单调缩放,而可靠性缩放是非线性的,4比特量化模型出现可靠性峰值,且量化增强了LLM对自然输入扰动的鲁棒性。

英文摘要

Quantization is a powerful strategy to build capable and resource-efficient large language models (LLMs) by reducing the bitwidth of the parameters. While quantized LLMs achieve state-of-the-art performance on unperturbed inputs using standard predictive metrics, their performance on perturbed inputs, measured using reliability metrics, remains underexplored, despite its importance for reliable deployment. To address this gap, we first conduct a comprehensive reliability evaluation of quantized LLMs consisting of three key components: (1) Uncertainty: We assess the trustworthiness of LLMs quantized to 2, 3, 4, and 8 bits using six different quantization methods, employing established uncertainty metrics. (2) Calibration: We assess how well-calibrated the uncertainty estimates of quantized models are across model scales and bit precisions. (3) Robustness: We design character-level and word-level input perturbations to evaluate the reliability of quantized models under semantically-preserving variations in the inputs that arise in real-world applications. Second, we characterize how reliability scales with the total number of model bits. Our study reveals that while the performance scales monotonically with the total number of bits, the reliability scalings are nonlinear. A reliability peak occurs for 4-bit quantized models, indicating that quantizing moderately sized models offers the best reliability-efficiency trade-off. Additionally, our empirical findings reveal that quantization enhances the robustness of LLMs to natural input perturbations.

CommentsTransactions on Machine Learning Research (TMLR), 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑