arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习恒等映射:SGD如何在函数分解中进行选择的案例研究

Learning the identity: a case study of how SGD selects among functional decompositions

Andy Arditi, Weian Xie, David Bau, Liu Ziyin

arXiv 2610.00615首次发表:更新:

发表机构

Northeastern University; Massachusetts Institute of Technology(东北大学; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以深度线性残差网络学习恒等函数为例,通过熵损失理论解析SGD在众多等价分解中偏好特定解的原因,并验证了理论预测。

AI 中文摘要

人们可能会认为,使用深度线性残差网络学习恒等函数是微不足道的——残差连接路径已经实现了恒等映射,因此网络只需将其权重驱动到零即可。然而,这个零权重解只是整个种群损失最小化器流形上的一个点,每个最小化器对应于恒等函数在网络各层之间的不同分解。尽管种群损失无法区分这些解,但随机梯度下降(SGD)可重复地偏好特定的解。例如,在非各向同性标签噪声下,学习到的层表现出依赖于噪声的谱;即使使用权重衰减,SGD通常也不会恢复零权重解。仅改变参数化方式,同时保持可实现函数集合不变,会产生不同的行为:将每个权重矩阵分解为两个矩阵的乘积,会导致权重坍缩到零,即使没有显式的权重衰减。虽然这些现象起初可能显得神秘且违反直觉,但可以通过熵损失(entropic loss)的视角来理解,该损失在种群损失基础上增加了一项与小批量梯度期望平方范数成比例的项(Ziyin等人,2025)。在恒等流形上,种群损失是常数,而熵项则区分了这些分解。我们解析地表征了其最小化器,并利用它们推导出SGD所偏好的解结构的预测。使用SGD训练的网络与这些预测高度吻合。总体而言,这里研究的恒等学习任务作为一个清晰而简单的案例研究,展示了熵损失的视角如何阐明SGD为何偏好同一输入-输出函数的特定分解。

英文摘要

One might think that learning the identity function with a deep linear residual network is trivial - the path along residual connections already implements the identity, and so the network need only drive its weights to zero. However, this zero-weight solution is just one point on an entire manifold of population-loss minimizers, each corresponding to a different decomposition of the identity across the network's layers. Although the population loss does not distinguish among these solutions, stochastic gradient descent (SGD) reproducibly favors particular ones. For instance, under anisotropic label noise, the learned layers exhibit a noise-dependent spectrum; even with weight decay, SGD does not generally recover the zero-weight solution. Changing only the parametrization, while leaving the set of realizable functions unchanged, yields different behavior: factoring each weight matrix as a product of two matrices causes the weights to collapse to zero, even without explicit weight decay. While perhaps mysterious and unintuitive at first, these phenomena can be understood through the lens of entropic loss, which augments the population loss with a term proportional to the expected squared norm of the minibatch gradient (Ziyin et al., 2025). On the identity manifold, the population loss is constant, while the entropic term distinguishes among these decompositions. We characterize its minimizers analytically and use them to derive predictions for the structure of solutions favored by SGD. Networks trained with SGD closely match these predictions. Overall, the identity learning task studied here serves as a clean and simple case study of how the lens of entropic loss can clarify why SGD favors particular decompositions of the same input-output function.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑