arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型(LLM)中量化损伤的结构:为何应将下一位精度预算用于全局分配

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

Jundong Hu, Shekar Ramachandran

arXiv 2609.01587首次发表:更新:

发表机构

PayPal AI(PayPal人工智能部门)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过因果混合精度干预探究LLM量化损伤的结构,发现精度恢复多为分散式,全局分配精度预算比局部修复关键层性能更优,需用因果干预验证量化损伤的定位。

AI 中文摘要

后训练量化(PTQ)被广泛用于降低大语言模型(LLM)推理服务的成本,但其精度损失分布不均,且通常需针对每个模型单独调优。本研究探究量化损伤发生的位置以及如何分配少量额外的精度预算:以因果混合精度干预为基准(依次将每层提升至8位,测量其恢复的精度),在4个架构族的9个开源权重模型上,我们测试了3个直观假设:量化损伤存在于任务回路、模型计算位置或权重统计中,但 none 能够预测哪些层会从恢复的精度中受益。精度恢复是分散的:在9个模型中的8个,恢复75%的精度缺口大致需要一半的层;唯一的例外是 Qwen3-8B,其精度恢复高度集中。在匹配的精度预算下,对于所有8个兼容group-128的模型(除OpenLLaMA外,其宽度不符合group-128要求),将预算全局用于更精细的量化粒度,比局部修复最具恢复潜力的层性能高出21至52个点,包括高度集中的Qwen3-8B。我们报告了2个次要发现:在RTN、GPTQ和AWQ的评估中,残差受预算限制(8位几乎无损失);峰值恢复位置在同一架构族内与架构相关,但在不同架构族间不相关。在该预算设置下,全局粒度是比选择性保护关键层更好的默认选择;更广泛地说,与量化损伤相关的廉价信号不一定能识别恢复精度可提升精度的位置,这必须通过因果干预进行测试。

英文摘要

Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground truth (raise each layer to 8-bit in turn and measure the accuracy it recovers) across 9 open-weight models in 4 architecture families, we test 3 intuitive hypotheses: that quantization damage lives in task circuits, where the model computes, or in weight statistics. None of them predicts which layers benefit from restored precision. Recovery is instead diffuse: for 8 of 9 models, recovering 75% of the gap takes roughly half the layers; the lone exception, Qwen3-8B, is sharply concentrated. At a matched precision budget, spending it globally on finer quantization granularity beats locally repairing the most recoverable layers for all 8 group-128-compatible models (all but OpenLLaMA, whose width rules out group-128), by 21-52 points, including the concentrated Qwen3-8B. We report 2 secondary findings: the residual is budget-limited (8-bit is near-lossless in our evaluation across RTN, GPTQ, and AWQ), and the location of peak recovery correlates with architecture within a family, though not across families. Within this budget setting, global granularity is a better default than selectively protecting critical layers. More broadly, cheap signals that correlate with quantization damage do not necessarily identify where restoring precision improves accuracy; this must be tested with causal intervention.

CommentsPreprint. Under review at a NeurIPS 2026 workshop. 11 pages, 4 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑