发表机构
National Central University; National Yang Ming Chiao Tung University(中央大学; 阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出将残差连接的收缩映射统一为单参数角收缩族,得到Proj-SpheretNorm、Cay-SpheretNorm等方法,在nanoGPT上验证其性能优于现有方案,发现指数映射并非球面残差流的最优选择。
AI 中文摘要
残差连接是稳定训练深度神经网络的实际机制。测地归一化(GeoNorm)将其重构到黎曼流形上,使每层输出与当前隐状态正交,并通过黎曼指数映射应用所得更新。因此,每个隐状态保持恒定的ℓ₂范数,将残差流限制在超球面上。然而,指数映射只是广泛收缩映射族中的一员。我们证明,在超球面上,整个该族坍缩为单个标量设计选择。区分不同收缩映射的仅在于更新的幅度如何转换为隐状态与更新所张平面内的旋转角。该观点将欧几里得残差连接与GeoNorm置于同一框架中。用度量投影收缩和Cayley收缩实例化它,得到Proj-SpheretNorm和Cay-SpheretNorm,二者均严格保范且仅需代数运算。两种方法均为单参数角收缩族p-SpheretNorm的成员,其旋转角达到饱和而非无界增长。上述两种方法分别在p=1和p=2时精确恢复,而恒等映射和GeoNorm仅在两端作为极限出现。在nanoGPT上,所有三种方法均优于现有轻量深度连接方案,且在有限p时取得最佳验证损失,表明指数映射并非球面残差流的优选收缩映射,而仅是该谱的一端。
英文摘要
Residual connections are the de facto mechanism for training deep neural networks stably. Geodesic Normalization (GeoNorm) recasts them on a Riemannian manifold, orthogonalizing each layer output against the current hidden state and applying the resulting update through the Riemannian exponential map. Every hidden state thus keeps a constant $\ell_{2}$-norm, confining the residual stream to a hypersphere. The exponential map, however, is only one member of a broad family of retraction maps. We show that on the hypersphere this entire family collapses to a single scalar design choice. What distinguishes one retraction from another is only how the magnitude of an update is converted into a rotation angle within the plane spanned by the hidden state and the update. This view places Euclidean residual connections and GeoNorm in one framework. Instantiating it with the metric projection retraction and the Cayley retraction yields Proj-SpheretNorm and Cay-SpheretNorm, which are exactly norm-preserving yet require only algebraic operations. Both prove to be members of a one-parameter family of angular retractions, $p$-SpheretNorm, whose rotation angle saturates rather than growing without bound. The two methods above are recovered exactly at $p = 1$ and $p = 2$, while the identity map and GeoNorm arise only as limits at either end. On nanoGPT, all three methods outperform existing lightweight deep connection schemes, and the best validation loss is attained at finite $p$, indicating that the exponential map is not the preferred retraction for spherical residual streams but merely one end of a spectrum.
Comments23 pages, 3 figures, and 6 tables