arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06685stat.MLcs.LG

RoPE注意力是保持softmax不变的精确前向传播梯度步

RoPE attention is an exact forward-pass gradient step with softmax intact

  • Hassana Labs(哈萨纳实验室)
  • Oxford University(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Julie Huang, Maggie Chlon, Leon Chlon

中文总结 AI 辅助

本文推导了RoPE-softmax注意力前向传播的精确梯度步表示,构造查询相关矩阵并利用指数差分保留softmax,验证了表示并量化重用误差。

中文摘要 AI 辅助

我们推导了RoPE-softmax前向传播的精确梯度步表示。对于任意具有仿射投影权重的确定性RoPE-softmax注意力头,我们构造了一个查询相关的有效矩阵$\Delta M_i$,满足$y_i = \mu_i + u_i^\top \Delta M_i$,其中$\mu_i$是被注意值的均匀均值,$u_i$是增广查询输入。该构造应用经典指数差分$\rho = \phi_1$来精确保留softmax。其正系数在查询条件下的二次目标上给出了单位梯度步表示。同一函数将RoPE生成器与精确位置有限差分联系起来。我们推导了重用单个查询矩阵的误差的逐词公式,并证明非恒定有限缓存头不能允许全局精确的仿射查询读出。在预训练的Qwen2.5-0.5B层上的重建检查和冻结重用校准验证了该表示,并量化了重用单个查询矩阵时所需的修正。

英文摘要

We derive an exact gradient-step representation of the RoPE-softmax forward pass. For every deterministic RoPE-softmax attention head with arbitrary affine projection weights, we construct a query-dependent effective matrix $ΔM_i$ satisfying $y_i = μ_i + u_i^\top ΔM_i$, where $μ_i$ is the uniform mean of the attended values and $u_i$ is the augmented query input. The construction applies the classical exponential divided difference $ρ= ϕ_1$ to retain the softmax exactly. Its positive coefficients give a unit gradient-step representation on a query-conditioned quadratic objective. The same function connects the RoPE generator to exact positional finite differences. We derive a tokenwise formula for the error of reusing one query's matrix and prove that a nonconstant finite-cache head cannot admit a globally exact affine query readout. Reconstruction checks and frozen-reuse calibration on one pretrained Qwen2.5-0.5B layer verify the representation and quantify the correction required when one query's matrix is reused.

↑