arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03640cs.LGq-bio.NCstat.ML

欠完备线性自编码器中的破缺尺度对称性

Broken scale symmetries in undercomplete linear autoencoders

Farhad Pashakhanloo, Jacob A. Zavatone-Veth

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示欠完备线性自编码器中SGD破坏尺度对称性,偏向大解码器权重,产生定向尺度漂移并受稳定性边界限制,展示损失几何将梯度噪声转为定向运动。

中文摘要 AI 辅助

神经网络损失景观具有许多对称性,这些对称性被梯度流所保持,但被有限步长的随机梯度下降(SGD)所破坏。此类对称性的一个典型例子是同质网络中的尺度对称性:可以放大某一层的参数并缩小下一层的参数,而不改变网络的输出。先前的工作已记录了SGD破坏这种对称性以偏向平衡梯度噪声或最小化波动的案例。在此,我们表明欠完备线性自编码器的解几何反而为尺度漂移选择了一个优先符号:在PCA解流形上,SGD倾向于大的解码器权重。这种定向的尺度漂移发生在慢时间尺度上,其动力学允许解析可处理的等效描述。然而,它不能无限期地持续:增大的尺度最终会将动力学推向有限步长的稳定性边界。在损失Hessian的最大特征值意义上,所得解比平衡基线更尖锐,但不同的尖锐度度量可能朝相反方向移动。因此,欠完备自编码器具体说明了损失几何如何将残余梯度噪声转化为沿功能等效解流形的定向运动。

英文摘要

Neural network loss landscapes have many symmetries, which are preserved by gradient flow but broken by finite-stepsize stochastic gradient descent (SGD). A canonical example of such a symmetry is scale in homogeneous networks: one can scale up the parameters in one layer and down in the next without changing the network output. Previous work has documented cases in which SGD breaks this symmetry in favor of balancing gradient noise or minimizing fluctuations. Here, we show that the solution geometry of undercomplete linear autoencoders instead selects a preferred sign for scale drift: on the PCA solution manifold, SGD favors large decoder weights. This directed scale drift occurs on a slow timescale, and its dynamics admit an analytically-tractable effective description. However, it cannot continue indefinitely: increasing scale eventually drives the dynamics towards a finite-stepsize stability boundary. The resulting solutions are sharper than a balanced baseline in the sense of the maximum eigenvalue of the loss Hessian, but different sharpness measures can move in opposing directions. Thus, undercomplete autoencoders give a concrete illustration of how loss geometry can convert residual gradient noise into directed motion along a manifold of functionally-equivalent solutions.

补充信息

↑