发表机构
StepFun; University of Chinese Academy of Sciences; Tsinghua University(阶跃星辰; 中国科学院大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HyperTransfer通过证明超球优化器与基础优化器在尺度不变网络中的动态等价性,仅用初始化和学习率调度即可复现目标优化器动态,并扩展至非尺度不变网络,实验验证了轨迹一致性。
AI 中文摘要
超球优化器约束参数范数并仅更新其方向,为神经网络优化建立了一种独特的范式。尽管这种几何结构与常规基础优化器的几何结构看似根本不同(后者同时更新参数范数和方向),但我们证明,对于尺度不变网络,这两种范式在动态上是等价的。基于这种等价性,我们提出了HyperTransfer,它仅利用目标基础优化器的初始化和学习率调度来构建一个能复现其动态的超球优化器,而无需运行目标优化器本身。我们进一步推导了逆映射,并将该框架扩展到非尺度不变网络。实验表明,HyperTransfer和逆映射产生的损失轨迹与其目标几乎相同,这表明超球动态主要由诱导的有效学习率调度和优化器状态所主导。
英文摘要
Hyperball optimizers constrain parameter norms and update only their directions, establishing a distinct paradigm for neural network optimization. Although this geometry appears fundamentally different from that of conventional Base Optimizers, which update both parameter norms and directions, we show that the two paradigms are dynamically equivalent for scale-invariant networks. Building on this equivalence, we propose HyperTransfer, which constructs a Hyperball optimizer that reproduces the dynamics of a target Base Optimizer using only its initialization and learning-rate schedule, without running the target optimizer itself. We further derive the inverse mapping and extend the framework to non-scale-invariant networks. Experiments show that both HyperTransfer and the inverse mapping produce loss trajectories nearly identical to those of their targets, suggesting that Hyperball dynamics are governed primarily by the induced effective learning-rate schedule and optimizer state.