从注意力敏感性到层角色:重新审视Transformer的混合精度量化
From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers
- Sharif University of Technology(谢里夫理工大学)
- GISMA University of Applied Sciences(GISMA应用科学大学)
- BRAINS, Brandenburg Research Center for Applied Intelligent Systems(勃兰登堡应用智能系统研究中心)
- Max Planck Institute for Security and Privacy(马克斯·普朗克安全与隐私研究所)
- Computer Vision Center (CVC)(计算机视觉中心)
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出JAB方法,基于注意力输出联合优化Q、K、V权重,并引入角色感知偏移规则,在Mistral-7B和GPT-2上实现高效混合精度量化,显著优于传统逐矩阵方法。
AI中文摘要:
大多数训练后量化流程将每个权重矩阵逐一拟合到其预训练对应物上。该代理是否追踪注意力块实际计算的内容,或者Q、K、V投影中的误差如何在softmax内部复合,很少被检查。我们改为在注意力输出上编写目标,同时覆盖所有三个投影,并在整个流程中重复使用它。JAB在块的联合Q、K、V权重上定义一个标量损失,根据块的真实因果掩码注意力输出进行评估,并使用两次:一次用于拟合量化权重(GPTQ热启动,然后使用可学习尺度的STE),另一次用于对块进行多选背包分配评分。在Mistral-7B的仅注意力量化中,这种方法有效。在3比特下,JAB恢复了均匀GPTQ与全精度之间差距的77-90%,其敏感性估计跟踪了需要73次前向传播的oracle,误差在几分之一分以内。一旦MLP层进入分配,它就停止工作。一个无需敏感性估计的角色感知偏移规则在GPT-2的MLP和完整的Mistral-7B模型上击败了JAB:在3比特下限下,它将96.4%的权重量化为每参数4.5比特,困惑度为6.933,在全精度(6.643)的4.4%以内,压缩比为3.56倍,而JAB在同一预算下为7.158。权重所在的矩阵比我们计算的任何敏感性估计都更重要。有两件事出乎意料。块局部重建是端到端困惑度的不可靠代理:一次运行将块自身目标提高了4.6倍,而困惑度上升了32倍,这就是为什么这里的每次分配都经过端到端验证。在仅注意力量化中,微调使权重远离其预训练值,同时将注意力输出拉近,带来净收益。训练后似乎恢复了注意力行为,而不是权重。
英文摘要:
Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objective on the attention output instead, over all three projections at once, and reuse it throughout the pipeline. JAB defines one scalar loss over the joint Q, K, V weights of a block, evaluated against the block's real causally-masked attention output, and uses it twice: to fit the quantized weights (GPTQ warm start, then STE with learnable scales), and to score the block for a multiple-choice knapsack allocation. On attention-only quantization of Mistral-7B this works. At 3 bits JAB recovers 77-90% of the gap between uniform GPTQ and full precision, and its sensitivity estimate tracks an oracle costing 73 forward passes to within a fraction of a point. It stops working once MLP layers enter the allocation. A role-aware offset rule needing no sensitivity estimate at all beats JAB on GPT-2's MLP and on the full Mistral-7B model: with a 3-bit floor it quantizes 96.4% of the weights to 4.5 bits per parameter at 6.933 perplexity, within 4.4% of full precision (6.643) at 3.56x compression, against 7.158 for JAB at the same budget. Which matrix a weight sits in matters more than any sensitivity estimate we computed. Two things came out sideways. Block-local reconstruction is an unreliable proxy for end-to-end perplexity: one run improved a block's own objective 4.6x while perplexity rose 32x, which is why every allocation here is validated end-to-end. And on attention-only quantization, fine-tuning moved weights farther from their pretrained values while pulling attention outputs closer, with net gains. Post-training seems to recover attention behavior, not weights.