当局部方差最优性不足时:用于动态4比特量化的RoPE对齐Q/K旋转
When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation
AI总结:
本文针对动态4比特量化,研究RoPE对齐的Q/K旋转,推导了相关旋转角度,发现结构化代理最优性在与量化器尺度统计量不匹配时无法降低量化误差。
AI中文摘要:
基于旋转的后训练量化通常会在整个注意力头应用正交变换,以减少异常值引发的误差。RoPE(旋转位置编码)则将每个注意力头划分为二维频率对,由此引出一个问题:遵循该分解的变换是否能优于全头混合。已有研究证实了与RoPE可交换的逐对旋转,本文则给出逆结果:对于不同频率,不存在其他单头正交映射能与RoPE交换。针对实验中采用的头共享参数化方式,本文推导了在池化协方差、位置平均代理下,使较大通道方差最小化的旋转角度,并验证了该实现达到其解析最小值。在测试的动态W4A4KV4设置下,所评估的头共享逐对配置未提升准确率;在四个检查点上,用该配置替换全头Hadamard变换会同时增加短、长上下文长度下的困惑度。将逐对旋转与Hadamard组合,在默认估计器下满足±0.05困惑度(PPL)区间准则。仅从K估计共享角度,在每个检查点上都优于仅逐对配置,但未弥合其与全头混合的差距。该解析目标控制池化校准协方差的位置平均二阶矩,而动态量化器则从逐词组范围设置其步长;逐对变换仅支持两通道混合。在从两通道到全头混合的可控插值中,K范围、相对量化误差和困惑度下降均随支持度增加而降低。这些结果表明,当代理与混合支持和量化器的尺度设定统计量不一致时,结构化代理的最优性未必能降低量化误差。
英文摘要:
Rotation-based post-training quantisation commonly applies an orthogonal transform across an entire attention head to reduce outlier-induced error. RoPE instead partitions each head into two-dimensional frequency pairs, raising the question of whether a transform respecting this decomposition can improve on full-head mixing. Prior work has established the per-pair rotations that commute with RoPE. We state the converse result that, for distinct frequencies, no other single-head orthogonal map commutes with RoPE. For the head-shared parameterisation used in our experiments, we then derive the rotation angle that minimises the larger channel variance under a pooled-covariance, position-averaged surrogate and verify that the implementation attains its analytic minimum. The evaluated head-shared pairwise configuration does not improve accuracy in the tested dynamic W4A4KV4 setting. Across four checkpoints, replacing the full-head Hadamard with this configuration increases perplexity at both short and long context lengths. Composing the pairwise rotation with the Hadamard satisfies the selected $\pm0.05$-PPL interval criterion under the default estimator. Estimating the shared angle from K alone improves pairwise-only on every checkpoint but does not close its gap to full-head mixing. The analytic objective controls a position-averaged second moment of a pooled calibration covariance, whereas the dynamic quantiser sets its step from a tokenwise group range. The pairwise transform also has only two-channel mixing support. Along a controlled interpolation from two-channel to full-head mixing, K range, relative quantisation error, and perplexity degradation decrease as support increases. These results show that optimality for a structured surrogate need not reduce quantisation error when the surrogate and mixing support are misaligned with the quantiser's scale-setting statistic.