arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Stiefel 注意力:当 Transformer 投影矩阵的几何决定优化器选择时——以及何时不决定

Directions That Don't Drift: Stiefel Manifold Routing for Transformer Attention

Rubén Darío Guerrero

arXiv 2609.19363首次发表:更新:

发表机构

NeuroTechNet S.A.S.(NeuroTechNet 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究将Transformer注意力投影矩阵约束到Stiefel流形并用黎曼Adam优化,发现其步长尺度自由特性带来显著性能提升,且权重衰减的零黎曼梯度使几何得以保持。

AI 中文摘要

注意力机制中的查询和键投影 $\WQ,\WK$ 几乎总是由欧几里得优化器训练,且对其几何形状没有约束。我们将它们约束到 Stiefel 流形上,并使用黎曼 Adam 优化器在该流形上进行优化,该优化器为每个帧携带一个标量二阶矩,通过信任区域限制其步长,并进行极坐标回缩。四个命题证明该更新在嵌入度量下是 steepest descent,与梯度尺度无关,条件良好,且精确地具有 $\mathrm{O}(d)$-等变性,每个命题均在 \texttt{float64} 下进行了数值验证。第五个命题提供了机制:权重衰减在 $\St(d,r)$ 上的黎曼梯度 \u003cemph{恒为零},因为 $W = W I_r$ 位于法空间中,因此学习到的注意力几何能够经受住衰减在模型其余部分驱动的崩溃循环。在模算术 grokking 任务上,单次运行在第 20,000 轮时保持 $97.0\\%$ 的验证准确率,而基线仅为 $61.1\\%$——我们将这个不稳定的终点作为该机制的证据而非效应量。在 CIFAR-10 patches 上,同样的规则在 12 对配对起始点上获得了 $\mathbf{+8.98}$ 个百分点($t{=}60.6$,$12/12$),且差距随数据增加而扩大而非缩小。步长规则带来了这一收益:固定步长的黎曼更新在梯度上是线性的,因此每步移动量比形状相同的 AdamW 矩阵少 $24$–$40$ 倍——其帧几乎不离开初始化位置,而完全冻结它们仅损失 $0.28$ 个百分点。消融实验将全部收益归因于使步长尺度自由,而投影器或等变性没有可测量的贡献。一个负面结果使该解释更加清晰:规范移除不能为该方法的动机提供依据,因为损失不变的方向上根本不携带梯度。

英文摘要

The query and key projections $\WQ,\WK$ in attention are almost always trained by Euclidean optimizers with no geometric constraint. We constrain them to the Stiefel manifold and optimize with a Riemannian Adam carrying one scalar second moment per frame---the form of \citet{becigneul2019}, here extended to the compact, non-Hadamard $\St(d,r)$ with a tangent projector, step-norm cap, and polar retraction. Four propositions prove steepest descent in the embedded metric, gradient-scale independence, well-conditioning, and exact $\mathrm{O}(d)$-equivariance. A fifth records that weight decay has \emph{identically zero} Riemannian gradient on $\St(d,r)$ ($W{=}WI_r$ lies in the normal space), so decay cannot act on the constrained frames. On a CIFAR-10 patch benchmark at $n{=}10\mathrm{k}$ this rule gains $\mathbf{+6.79}$\,pp over AdamW across 12 paired starts ($t{=}38.33$, $12/12$); earlier fixed-step Riemannian SGD gains $+1.97$\,pp, of which $+1.69$\,pp comes from frozen orthonormal initialization alone. The corrected Adam's lead grows with data: $+1.9$\,pp at $n{=}1\mathrm{k}$ to $+6.7$\,pp at $n{=}50\mathrm{k}$. A 12-seed ablation credits all gain to the scale-free step ($+4.63$\,pp, $12/12$), nothing to the projector or equivariance; a targeted $\varepsilon$-sweep causally confirms the mechanism ($-2.6$\,pp at $\varepsilon{=}0.1$, $p{<}0.001$). Two five-seed grokking studies confirm the constrained arm does not grok better than the baseline ($p{=}0.019$, A2 wins): the weight-decay exemption has no grokking consequence. A single-seed pilot exploiting this localization achieves the first stable grokking under slingshot conditions---Stiefel + targeted circuit regularization keeps routing-frame isometry error $10^6\times$ lower than the unconstrained ablation through every collapse.

Comments26 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑