发表机构
College of Pharmacy, Chungnam National University; Department of Computer Science & Engineering, Chungnam National University(忠南国立大学药学院; 忠南国立大学计算机科学与工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对均方根归一化问题,提出MRSNorm方法,通过将通道配对成相量改变缩放范式,约束激活在相量流形,共享权重减半参数,经分析和实验验证其能确保梯度均匀性、提供结构稳定性及防止数值爆炸,推动向基于相量的深度表示学习转变。
AI 中文摘要
均方根归一化已成为加速现代序列模型的事实上的标准,但其对独立标量二次积累($\sum x^2$)的依赖会引发异常值引起的数值不稳定、梯度饥饿和各向异性相位失真。我们引入均方根归一化(MRSNorm)。通过将通道结构配对成二维相量,MRSNorm在数学上反转了传统缩放范式:先计算局部$L_2$幅度(均方根),再通过全局$L_1$平均(均值)聚合。这种操作反转将激活严格约束在相量流形上,保持共形不变性。通过在相量分量间共享单个仿射权重,MRSNorm将可学习参数总数减半。分析表明这种几何约束产生由勾股定理控制的内置三角梯度裁剪器,确保梯度均匀性。在CIFAR - 100的ResNet上的实证评估表明,尽管参数减半,MRSNorm在严格压力测试下提供关键结构稳定性,在标准归一化遭受梯度发散的极端超参数设置下,成功防止数值爆炸并确保稳定优化轨迹。我们的发现提出了向基于相量的深度表示学习的基本范式转变。MRSNorm的实现见附录C。
英文摘要
While Root Mean Square Normalization has become the de facto standard for accelerating modern sequence models, its reliance on the quadratic accumulation of independent scalars ($\sum x^2$) inherently triggers outlier-induced numerical instability, gradient starvation, and anisotropic phase distortion. We introduce Mean Root Square Normalization (MRSNorm). By structurally pairing channels into 2D phasors, MRSNorm mathematically inverts the traditional scaling paradigm: it computes the localized $L_2$ magnitudes (Root Square) before aggregating them via a global $L_1$ average (Mean). This operational inversion strictly constrains activations to a phasor manifold, preserving conformal invariance. By sharing a single affine weight across phasor components, MRSNorm halves the total number of learnable parameters, proving that unconstrained spatial scaling in standard norms is a harmful redundancy. We analytically demonstrate that this geometric constraint yields a built-in, trigonometric gradient clipper governed by the Pythagorean identity, unconditionally equalizing the local gradient norm to ensure Gradient Homogeneity. Empirical evaluations on a ResNet with CIFAR-100 show that despite halved parameters, MRSNorm provides critical structural stability under rigorous stress tests. Under extreme hyperparameter settings where standard normalizations suffer from gradient divergence, MRSNorm successfully prevents numerical explosion and secures stable optimization trajectories. Our findings propose a fundamental paradigm shift toward phasor-based deep representation learning. The implementation of MRSNorm is available at Appendix C.