arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Transformer优化中的隐藏边界运动:仿射权重和偏差更新的函数空间正交化

Hidden Boundary Motion in Transformer Optimization: Function-Space Orthogonalization of Affine Weight and Bias Updates

Zhang Gongyue, Sheng Yixuan, Liu donghan, Wang Zhiyong, Ren Weihong, Liu honghai

arXiv 2607.22927首次发表:更新:

AI 中文总结

研究Transformer优化中权重与偏差更新的隐藏边界运动,提出将更新分解为形状和边界分量,实现SBO-AdamW优化器,实验表明其提升了准确率,但存在问题,支持隐藏边界运动为重要机制,下一步需稳定的中心仿射参数化。

AI 中文摘要

权重和偏差通常作为单独的参数张量进行优化,但当仿射层输入均值不为零时,它们并不代表独立的函数。对于仿射映射\(z = Wx + b\),权重更新包含与样本无关的位移\(\Delta W\mu\),在功能上与偏差更新无法区分,我们称这种隐藏贡献为“边界运动”,并将每次更新分解为中心的、样本变化的“形状”分量和共享的“边界”分量。在从零开始训练的四层Transformer上,偏差项\(g_b\mu^\top\)的中位数范数等于仿射层和训练检查点上原始权重梯度范数的0.664。更惊人的是,中位数比率\(\norm{\Delta W\mu}/\norm{\Delta b}\)为134.7,而\(\norm{\Delta W\mu}/\norm{\Delta b + \Delta W\mu}\)为0.994。因此,在AdamW下,观察到的边界运动几乎完全通过权重矩阵而非显式偏差实现。我们实现了诊断优化器Shape--Boundary Orthogonal AdamW (SBO-AdamW),它使用独立的Adam状态优化\(g_W - g_b\mu^\top\)和\(g_b\),并补偿权重引起的边界位移。在单种子实验中,SBO-AdamW将验证准确率从81.68%提高到85.81%,验证选择的测试准确率从78.73%提高到82.73%,最佳验证检查点出现在第800步而非第3000步。然而,移动批次中心补偿会产生严重的偏差坐标漂移并强烈降低边界能量。因此,现有证据支持隐藏边界运动作为一种重要的优化机制,但尚未建立最终的通用优化器。稳定的中心仿射参数化被确定为下一步所需的步骤。

英文摘要

Weights and biases are normally optimized as separate parameter tensors, yet they do not represent separate functions when the input to an affine layer has nonzero mean. For an affine map $z=Wx+b$ with input mean $μ$, a weight update contains a sample-independent displacement $ΔWμ$ that is functionally indistinguishable from a bias update. We call this hidden contribution \emph{boundary motion} and decompose each update into a centered, sample-varying \emph{shape} component and a shared \emph{boundary} component. On a four-layer Transformer trained from scratch on IMDb, the bias-like term $g_bμ^\top$ has a median norm equal to 0.664 of the raw weight-gradient norm across affine layers and training checkpoints. More strikingly, the median ratio $\norm{ΔWμ}/\norm{Δb}$ is 134.7, while $\norm{ΔWμ}/\norm{Δb+ΔWμ}$ is 0.994. Thus, under AdamW, the observed boundary motion is almost entirely realized through the weight matrix rather than the explicit bias. We implement a diagnostic optimizer, Shape--Boundary Orthogonal AdamW (SBO-AdamW), that optimizes $g_W-g_bμ^\top$ and $g_b$ with independent Adam states and compensates the weight-induced boundary displacement. In a single-seed experiment, SBO-AdamW raises validation accuracy from 81.68\% to 85.81\% and validation-selected test accuracy from 78.73\% to 82.73\%, with the best validation checkpoint occurring at step 800 instead of step 3000. However, the moving-batch-center compensation produces severe bias-coordinate drift and strongly reduces boundary energy. The present evidence therefore supports hidden boundary motion as an important optimization mechanism, but it does not yet establish a final general-purpose optimizer. A stable centered-affine parameterization is identified as the required next step.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑