AI 中文总结
研究无穷小初始化下对角线性网络梯度流动力学,扩展相关定理分析,提出算法1等效表征训练轨迹,证明其收敛到修改后的\(\mathcal{l}_1\)范数最小化问题解,确定隐式偏差并揭示关键几何结构。
AI 中文摘要
我们研究了无穷小初始化下用于回归任务的对角线性网络的梯度流动力学。扩展了Pesme和Flammarion(2023)的定理1,我们将分析推广到深度对角线性网络和更广泛的一类两层对角线性网络(如定义4.1中所定义)。具体而言,我们证明这些模型的训练轨迹可以由所提出的算法1等效地表征。我们进一步证明该算法收敛到一个修改后的\(\mathcal{l}_1\)范数最小化问题的解。结果,我们确定在无穷小初始化的情况下,两种网络架构的隐式偏差都对应于一个修改后的\(\mathcal{l}_1\)范数。此外,我们通过将结构不变流形(SIM)(Zhao等人,2026)识别为塑造学习过程的关键几何结构,深入了解了控制这些动力学的潜在机制。
英文摘要
We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as defined in Definition 4.1). Specifically, we demonstrate that the training trajectories of these models can be equivalently characterized by the proposed Algorithm 1. We further prove that this algorithm converges to the solution of a modified $ \mathcal{l}_1 $ norm minimization problem. As a result, we establish that the implicit bias of both network architectures corresponds to a modified $ \mathcal{l}_1 $ norm in the regime of infinitesimal initialization. Additionally, we provide insights into the underlying mechanisms governing these dynamics by identifying the Structural Invariant Manifold (SIM) (Zhao et al., 2026) as the key geometric structure that shapes the learning process.