量化循环Transformer:反馈暴露与校准盲区
Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
浏览论文内容
中文总结 AI 辅助
研究循环Transformer量化中的两种失效模式:反馈暴露与校准盲区,提出跨步骤累积GPTQ Hessian方法,在九个检查点上优于基线并恢复bf16精度。
中文摘要 AI 辅助
循环Transformer在循环步骤间复用权重,使得低位量化尤为具有吸引力。我们识别出标准训练后量化的两种不同失效模式。在Huginn-3.5B上,按通道INT4量化主要失败于非残差循环入口适配器,而对残差核心进行量化则破坏性小得多。我们称之为反馈暴露:量化层扰动了循环状态而无恒等路径,由此产生的误差在后续步骤中被反馈。在线性滤波器和Mamba状态空间模型上的受控实验表明,反馈暴露也发生在Transformer之外。分组INT4揭示了另一种失效,即校准盲区:我们的单步GPTQ基线从第0步激活构建其Hessian矩阵,使得循环中后续使用的输入方向几乎未被加权。在来自七个循环架构的九个检查点上,单步GPTQ在五个检查点上的主任务指标上差于最近舍入(RTN)。跨循环步骤累积GPTQ Hessian在所有九个检查点上均优于单步GPTQ和RTN,并在Huginn上恢复了bf16级精度。这些结果将循环模型上PTQ的两个问题区分开来:量化误差在何处进入循环,以及校准看到哪些状态。
英文摘要
Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual loop-entry adapter, while quantizing the residual core is much less damaging. We call this feedback exposure: a quantized layer perturbs the recurrent state without an identity path, and the resulting error is fed back at later steps. Controlled experiments on linear filters and Mamba state-space models show that feedback exposure also occurs outside transformers. Grouped INT4 reveals a separate failure, calibration blindness: our one-step GPTQ baseline builds its Hessian from step-0 activations, leaving input directions used later in the recurrence nearly unweighted. Across nine checkpoints from seven looped architectures, one-step GPTQ is worse than round-to-nearest (RTN) on the primary task metric for five checkpoints. Accumulating the GPTQ Hessian across recurrence steps outperforms both one-step GPTQ and RTN on all nine checkpoints and recovers bf16-level accuracy on Huginn. These results separate two questions for PTQ on looped models: where quantization error enters the recurrence, and which states calibration sees.
发表机构
- Meta
机构由 AI 辅助整理,请以论文原文为准。