arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

组件类型,而非重建误差,预测注意力量化的敏感性

Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity

Kasun Dewage, Marianna Pensky, Suranadi De Silva

arXiv 2609.26173首次发表:更新:

AI 中文总结

本研究通过大规模实验发现,注意力投影的组件类型(尤其是V投影)比重建误差更能预测量化敏感性,为混合精度量化位分配提供指导。

AI 中文摘要

许多训练后量化(PTQ)方法使用逐层重建、二阶代理目标或激活感知变换来减少量化引起的误差。然而,该误差信号是否能预测量化单个注意力投影的下游功能影响,尚未被直接表征。我们扫描了九个开放权重语言模型(1.3B--8B参数;OPT、GPT-J、LLaMA-1/2/3、Mistral、Qwen 2.5),并在最近舍入(RTN)下以及对于七个模型在GPTQ下,以3位和4位精度一次量化一个注意力投影,记录了3,808个不同测量中的重建误差、困惑度变化和每投影激活加权量化误差。我们发现:(1)在给定组件类型(Q、K、V或O)内,重建误差在RTN下的36种情况中的27种中解释了少于10%的困惑度敏感性方差,中位R²=0.044;(2)在所有9个模型中,组件类型和层身份都比重建误差解释更多方差,其中层身份在9个模型中的7个中是最强预测因子,组件类型在其余2个中最强;(3)值(V)投影是最常见的优势组件,在九个模型中的七个中占总正Delta PPL的38--51%;(4)优势组件在RTN和GPTQ之间大致保持(7例中的5例);(5)对于V投影,激活加权量化误差在组件内预测方面比重建误差适度更好(中位R²为0.20对0.06)。这些发现表明,仅相对权重重建误差不足以进行敏感性感知的位分配,并且V投影在混合精度方案中值得特别考虑。

英文摘要

Many post-training quantization (PTQ) methods use layer-wise reconstruction, second-order proxy objectives, or activation-aware transformations to reduce quantization-induced error. Whether that error signal predicts the downstream functional impact of quantizing an individual attention projection has not been directly characterized. We sweep nine open-weight language models (1.3B--8B parameters; OPT, GPT-J, LLaMA-1/2/3, Mistral, Qwen 2.5) and quantize one attention projection at a time under round-to-nearest (RTN) and, for seven models, GPTQ at 3 and 4 bits, recording reconstruction error, perplexity change, and per-projection activation-weighted quantization error for 3,808 distinct measurements. We find: (1) within a given component type (Q, K, V, or O), reconstruction error explains less than 10% of the variance in perplexity sensitivity in 27 of 36 cases under RTN, with median R^2 = 0.044; (2) both component type and layer identity explain more variance than reconstruction error in all 9 models, with layer identity the strongest predictor in 7 of 9 models and component type strongest in the remaining 2; (3) value (V) projections are the most commonly dominant component, accounting for 38--51% of total positive Delta PPL in seven of nine models; (4) the dominant component is broadly preserved between RTN and GPTQ (5 of 7 cases); and (5) activation-weighted quantization error is a moderately better within-component predictor than reconstruction error for V projections specifically (median R^2 of 0.20 vs. 0.06). These findings indicate that relative weight reconstruction error alone is insufficient for sensitivity-aware bit allocation, and that V projections merit dedicated consideration in mixed-precision schemes.

CommentsAccepted as a regular paper at IEEE ICMLA 2026; to appear in the conference proceedings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑