发表机构
Université Paris Dauphine-PSL; LORIA CNRS; ESPCI-PSL(巴黎多芬大学-PSL大学; 法国国家科学研究中心LORIA实验室; 巴黎市政工程工业物理和化学高等学校-PSL大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究将增长视为渐进式约束放松,提出迭代扩展可训练参数的策略,证明其可使训练偏向平坦区域,虽能生成平坦解但曲率降低未必提升测试性能,揭示平坦性与泛化关联的微妙性。
AI 中文摘要
深度神经网络尽管具有高度非凸、过参数化的损失景观,仍能很好地泛化,这一现象常与随机优化找到的极小值几何相关。我们将增长视为渐进式约束放松,研究增量式增长-优化策略如何使训练偏向更平坦的区域。从低维子模型出发,我们通过解锁嵌套随机子空间、冻结网络初始化时的正交补,迭代扩展可训练参数,每次扩展后重新优化,直至达到完整架构。在非退化极小值周围的标准局部正则条件下,我们证明局部次水平集可被椭球良好近似,且冻结约束下的盆地可达性可由冻结方向上的显式有效曲率表征。这解释了该策略的偏向性:渐进式增长通过冻结约束诱导的体积效应,增加了宽盆地的相对权重并抑制了尖锐盆地。我们在受控玩具景观和真实的ResNet/CIFAR-100设置中实证验证了这些预测,确认尽管渐进式子空间增长能可靠生成更平坦的解,但曲率降低并不普遍转化为测试性能提升,凸显了平坦性-泛化关联中的微妙之处。代码可在此httpsURL获取。
英文摘要
Deep neural networks generalize well despite their highly nonconvex, overparameterized loss landscapes, a phenomenon often associated with the geometry of the minima found by stochastic optimization. We study how incremental grow-and-optimize strategies bias training toward flatter regions by viewing growth as progressive constraint relaxation. Starting from a low-dimensional submodel, we iteratively expand the trainable parameters by unlocking nested random subspaces while freezing the orthogonal complement at the network initialization, re-optimizing after each expansion until the full architecture is reached. Under standard local regularity conditions around non-degenerate minima, we prove that local sublevel sets are well approximated by ellipsoids and that basin accessibility under frozen constraints can be characterized by an explicit effective curvature in the frozen directions. This leads to an explanation of the bias: progressive growth increases the relative weight of wide basins and suppresses sharp ones through a volume effect induced by the frozen constraints. We empirically validate these predictions in controlled toy landscapes and in a realistic ResNet/CIFAR-100 setting and confirm that although progressive subspace growth reliably produces flatter solutions, curvature reductions do not universally translate into improved test performance, highlighting subtleties in the flatness-generalization connection. The code is available at https://github.com/p0lcAi/Across-the-Loss-Landscape.
CommentsAccepted at ECML PKDD 2026