发表机构
The University of Tokyo; RIKEN Center for Advanced Intelligence Project; Kyoto University(东京大学; 理化学研究所先进智能研究中心; 京都大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究浅层线性网络在训练时域随宽度增长时的学习率快速迁移,证明在特定谱假设下迁移成立,并刻画了迁移速率及极限分布。
AI 中文摘要
跨模型宽度的超参数迁移可以大幅降低大型神经网络调参的成本,但其在训练时域随宽度增长时的行为尚未被完全理解。基于快速超参数迁移(Ghosh 等人,2026)的框架,该框架形式化了迁移何时有效,我们研究了在增长时域机制下确保快速迁移的条件。具体而言,我们研究了一个具有单个可训练隐藏矩阵的浅层线性网络中的学习率迁移,该网络通过全批量梯度下降进行训练。在额外的谱假设下,我们的主要结果有三方面:(i)我们证明了当 $n,T\to\infty$ 且 $T=o(\sqrt{n})$ 时,学习率快速迁移成立。(ii)我们通过有限宽度扰动尺度、损失及其学习率导数对有限宽度扰动的一阶敏感性以及局部损失曲率来刻画迁移速率。(iii)我们推导了最优学习率和优化损失的极限分布,这些分布由数据格拉姆矩阵的极端特征值相关的波动所控制。这些结果阐明了谱结构和局部损失敏感性如何在增长时域下主导学习率迁移。
英文摘要
Hyperparameter transfer across model width can substantially reduce the cost of tuning large neural networks, but its behavior when the training horizon grows with width is not fully understood. Building on the framework of fast hyperparameter transfer (Ghosh et al., 2026), which formalizes when transfer is effective, we investigate conditions that ensure fast transfer in the growing-horizon regime. Specifically, we study learning-rate transfer in a shallow linear network with a single trainable hidden matrix, trained by full-batch gradient descent. Under additional spectral assumptions, our main results are threefold. (i) We prove fast learning-rate transfer as $n,T\to\infty$ whenever $T=o(\sqrt{n})$. (ii) We characterize the transfer rates through the finite-width perturbation scale, the first-order sensitivities of the loss and its learning-rate derivative to finite-width perturbations, and the local loss curvature. (iii) We derive limiting distributions for the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix. These results clarify how spectral structure and local loss sensitivities govern learning-rate transfer at growing horizons.