arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02373cs.LGcond-mat.dis-nncond-mat.stat-mechcs.AI

优化中的渗流动力学:方差级联与离散标度不变性

Percolation Dynamics in Optimization: Variance Cascades and Nested Symmetry

发表机构慕尼黑工业大学计算、信息与技术学院 · 慕尼黑机器学习中心
查看机构详情
  • School of Computation, Information and Technology, Technical University of Munich(慕尼黑工业大学计算、信息与技术学院)
  • Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

Sai Niranjan Ramachandran, Suvrit Sra

首次发表
浏览论文内容

中文总结 AI 辅助

该研究将随机梯度流建模为渗流过程,揭示SGD引导深度神经网络子网络合并的离散同步块机制,且该机制可扩展至Adam和AdamW,为优化动力学提供了新的物理解释。

中文摘要 AI 辅助

我们研究随机梯度下降(SGD)的动力学,已知该算法会引导深度神经网络趋向对应更简单子网络的不变集,但这种引导随时间的展开过程仍知之甚少。我们通过将随机梯度流(SGF)建模为渗流过程来解答这一问题,其中架构对称性迫使子网络以离散的同步块合并,而非逐个合并。这些结构转变表现为宏观序参量中的方差尖峰,与物理相变呼应。我们进一步证明,在显式重尾噪声模型下,该陷阱机制及其相关的标度级联可扩展至Adam和AdamW。

英文摘要

We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which nested architectural symmetries force subnetworks to merge in discrete blocks rather than by single-edge attachment. These structural transitions register as variance spikes in a macroscopic order parameter echoing physical phase transitions. We further state sufficient conditions under which the trapping argument carries over to Adam and AdamW under heavy-tailed gradient noise and measure them on a trained Transformer.

补充信息

↑