发表机构
University of California, Davis(加州大学戴维斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出均匀相位初始化,利用正弦函数的周期对称性避免中心极限定理的近似误差并解耦层间依赖,在神经表示任务上超越现有最先进方法,且支持宽度扩展。
AI 中文摘要
深度神经网络的成功训练高度依赖于初始权重的分布。如果权重过大,网络训练会发散;如果过小,模型则无法学习特征。稳定初始化是这两个极端之间的最优折中。传统的随机网络理论使用中心极限定理来控制神经元间的依赖性,这引入了分布近似误差和层间耦合。对于使用正弦激活函数的网络,我们推导出了均匀相位初始化,它避免了分布近似并完全解耦各层。我们是首个利用正弦函数周期对称性的工作。使用均匀相位初始化训练的模型在图像和音频拟合等神经表示任务中优于现有最先进方法。我们发现,我们未经调优的模型与以往工作中最佳调优的基线相当,并支持$\mu$P宽度扩展。
英文摘要
Successful training of deep neural networks is highly dependent on the distribution of the initial weights. If the weights are too large, network training blows up; if they are too small, the model fails to learn features. Stable initialization is the optimal moderation between these two extremes. The conventional theory of random networks uses the Central Limit Theorem to control inter-neuron dependencies, which introduces distributional approximation error and coupling between layers. For networks with sine activations, we derive the uniform-phase initialization, which obviates distributional approximation and fully decouples the layers. Ours is the first work to use the sine function's periodic symmetry. Models trained with the uniform-phase initialization outperform the state of the art in neural representation tasks like image and audio fitting. We find that our untuned models are competitive with the best-tuned baselines from previous work and support $μ$P width scaling.
Commentsreproduction code in ancillary material