发表机构
École Polytechnique Fédérale de Lausanne; Institute of Mathematics; Department of Mathematics(联邦理工学院洛桑校区; 数学研究所; 数学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究随机失活与随机梯度掩码两种训练技术,发现在大型残差网络中,二者在深度和宽度渐近大时差异消失,都收敛到相同极限动态,此渐近等价性适用于多种变体,包括逐层随机失活。
AI 中文摘要
随机失活(Dropout)和随机梯度掩码(RaM)是深度学习中用于提高性能的两种训练技术。二者都将随机性引入训练动态,但方式不同。随机失活在前向传播中对激活值应用随机掩码,而RaM保持前向传播不变,而是对梯度进行掩码。特别是,RaM在参数更新中引入的噪声是无偏的,所以随机失活有效性的标准解释不适用于RaM。本文表明,在深度和宽度渐近大的情况下,对于残差网络,两种方法的差异消失:在完全特征学习 regime中,它们都收敛到相同的大规模极限动态。这种渐近等价性适用于随机失活和RaM的几种变体,包括随机深度残差网络中使用的逐层随机失活,尽管定量速率较慢。事实上,我们还表明,其中一些变体渐近地收敛到相同的极限。
英文摘要
Dropout and Random Gradient Masking (RaM) are two training techniques used to improve performance in deep learning. Both techniques inject randomness into the training dynamics, but in significantly different ways: dropout applies random masks to the activations in the forward pass, whereas RaM leaves the forward pass unchanged and instead masks the gradients. In particular, the noise induced by RaM in the parameter updates is unbiased, so standard explanations for the effectiveness of dropout, such as the penalization effect or the prevention of co-adaptation between neurons, do not apply to RaM. In this work, we show that the difference between the two methods disappears for ResNets in the large depth and width asymptotics: in the complete feature learning regime, they both converge to the same large-scale limiting dynamics. This asymptotic equivalence holds for several variants of dropout and RaM, including layerwise dropout as used in stochastic-depth ResNets, albeit at slower quantitative rates. In fact, we also show that several of these variants collapse to the same limit asymptotically.