arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习正交多指标模型超越小初始化:增量学习、竞争动力学与对称性

Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry

Mo Zhou, Weihang Xu, Simon S. Du, Maryam Fazel

arXiv 2609.10879首次发表:更新:

发表机构

University of Washington; Amazon, Inc.(华盛顿大学; 亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究证明标准初始化下多项式宽度两层网络学习正交多指标目标时仍发生增量学习,并揭示参数质量的竞争性再分配动力学,通过对称化网络近似实现理论分析。

AI 中文摘要

近期工作已识别出在单指标和多指标模型上训练的浅层网络中的增量学习现象。然而,现有分析往往依赖于简化设置,如小初始化、相关损失或逐层训练。这些选择减少了神经元间的相互作用,并使得标准初始化下的一些特征学习动力学未被探索。我们研究了在标准初始化下,使用多项式数量样本训练多项式宽度两层网络学习正交多指标目标时的训练动力学。我们首先证明增量学习仍然发生:损失根据目标的Hermite展开依次下降,低阶分量先于高阶分量被学习,并恢复各个目标方向。在这一标准初始化机制下,训练还表现出参数质量的竞争性再分配:在总质量拟合目标均值并稳定后,质量转移到目标子空间,然后集中于对齐的神经元上。我们的理论分析使用了略微修改的梯度流,而普通梯度下降在经验上表现出相同的定性动力学。在技术上,我们通过对称化网络引入了一种基于对称性的有限宽度近似,而非直接与无限宽度极限比较。这能更好地控制近似误差,并可能具有独立的研究价值。

英文摘要

Recent work has identified incremental learning in shallow networks trained on single-index and multi-index models. However, existing analyses often rely on simplifying settings, such as small initialization, correlation loss, or layer-wise training. These choices reduce neuron interactions and leave some feature learning dynamics under standard initialization unexplored. We study training dynamics for polynomial-width two-layer networks learning orthogonal multi-index targets under standard initialization using polynomially many samples. We first prove that incremental learning still occurs: the loss decreases sequentially according to the Hermite expansion of the target, with lower-order components learned before higher-order components recover the individual target directions. In this standard initialization regime, training also shows a competitive reallocation of parameter mass: after the total mass fits the target mean and stabilizes, mass shifts into the target subspace and then concentrates on aligned neurons. Our theoretical analysis uses slightly modified gradient flow, while vanilla gradient descent empirically exhibits the same qualitative dynamics. Technically, we introduce a symmetry-based finite-width approximation via symmetrized networks, rather than comparing directly with an infinite-width limit. This yields better control of approximation errors and may be of independent interest.

Comments102 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑