发表机构
Canva Research(Canva研究公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对神经网络操作乘法性质及现有优化器问题,提出指数线性权重重参数化方法\method,结合对称指数与线性路径,经实验验证其能提升损失下降速度,在变压器训练中减少训练步骤,还有不匹配初始化可改善早期优化。
AI 中文摘要
许多神经网络操作具有乘法性质而非加法性质。像Adam这样的自适应优化器按坐标归一化更新,但更新步骤仍是加法性的。我们引入了\textbf{\method}(\textbf{\methodshort}),一种神经网络权重重参数化方法,它将符号感知对称指数路径与恒等线性路径相结合。对称指数路径在小原始权重时近似线性,在大权重时曲率增加。对数空间中的加法更新映射到有效权重空间中与幅度成比例的变化。线性路径提供了稳定优化的直接途径,可学习的尺度、曲率和偏移参数控制路径间平衡及指数路径曲率。实验表明,这种方法在经验上提高了损失下降速度。我们还确定了一种有用的\emph{不匹配初始化},在小模型消融实验中,这改善了早期优化。我们在OpenWebText上对九种宽度×深度配置的变压器进行训练,\methodshort在少1.32 - 1.49倍训练步骤下达到匹配验证损失,宽度越大收益越大。
英文摘要
Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps. Adaptive optimizers such as Adam normalize updates per coordinate, but update steps remain additive; weights with very different magnitudes receive similarly sized absolute changes, producing very different relative perturbations. We introduce \textbf{\method} (\textbf{\methodshort}), a weight reparameterization for neural networks that combines a sign-aware symmetric-exponential pathway with an identity-like linear pathway. The symmetric-exponential pathway is near-linear for small raw weights but increasingly curved at larger magnitudes. Additive updates in logarithmic space map to magnitude-proportional changes in effective weight space. The linear pathway provides a direct route through the transform that we hypothesize stabilizes optimization, while learnable scale, curvature, and offset parameters control balance between pathways and the curvature of the exponential pathway. These components create a curved parameter-space geometry that empirically improves speed of loss descent over standard linear parameterization. We also identify a useful \emph{mismatched initialization}: raw weights are chosen so a symmetric version of the transform matches Xavier statistics, but training uses an asymmetric forward transform that leaves positive weights at full strength while making negative weights smaller in magnitude; in small-model ablations, this improves early optimization and may act as a form of symmetry breaking. We train transformers on OpenWebText over nine width$\times$depth configurations, \methodshort reaches matched validation loss in 1.32--1.49$\times$ fewer training steps, with the largest widths seeing the biggest gains.
Comments27 pages, 16 figures