优化中的渗流动力学:方差级联与离散标度不变性
Percolation Dynamics in Optimization: Variance Cascades and Nested Symmetry
查看机构详情
- School of Computation, Information and Technology, Technical University of Munich(慕尼黑工业大学计算、信息与技术学院)
- Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究将随机梯度流建模为渗流过程,揭示SGD引导深度神经网络子网络合并的离散同步块机制,且该机制可扩展至Adam和AdamW,为优化动力学提供了新的物理解释。
中文摘要 AI 辅助
我们研究随机梯度下降(SGD)的动力学,已知该算法会引导深度神经网络趋向对应更简单子网络的不变集,但这种引导随时间的展开过程仍知之甚少。我们通过将随机梯度流(SGF)建模为渗流过程来解答这一问题,其中架构对称性迫使子网络以离散的同步块合并,而非逐个合并。这些结构转变表现为宏观序参量中的方差尖峰,与物理相变呼应。我们进一步证明,在显式重尾噪声模型下,该陷阱机制及其相关的标度级联可扩展至Adam和AdamW。
英文摘要
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which nested architectural symmetries force subnetworks to merge in discrete blocks rather than by single-edge attachment. These structural transitions register as variance spikes in a macroscopic order parameter echoing physical phase transitions. We further state sufficient conditions under which the trapping argument carries over to Adam and AdamW under heavy-tailed gradient noise and measure them on a trained Transformer.