arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超球可能并非免费午餐

Hyperball May Not Be a Free Lunch

Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai

arXiv 2607.22444首次发表:更新:

发表机构

IQuest Research; Peking University; Sun Yat-sen University; Shenzhen University of Advanced Technology; Shanghai University of Finance and Economics(IQuest研究公司; 北京大学; 中山大学; 深圳先进技术大学; 上海财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究超球风格优化器优势来源,通过推导角有效学习率等方法分析,发现其优势并非源于更新方向,而是有效步长演变,预训练实验表明学习率衰减策略影响性能,谨慎调度对发挥其潜力至关重要。

AI 中文摘要

对于尺度不变的深度网络,超球风格的优化器通过固定矩阵值参数的范数和归一化更新在大规模训练中表现出强大性能。但其优势来源不明。本文从连续参数状态间的角位移出发,推导了角有效学习率,表明传统基于范数的度量是参数更新正交性下的特殊情况。分解优化器更新为径向和切向分量,分析径向更新对角位移的影响。数值结果显示径向分量对角有效学习率的直接影响有限,无法解释MuonH在训练初期比MuonWD收敛慢但后期超过它的原因。通过启发式实验发现它们的主要差异源于有效步长的演变而非超球诱导的本质上更优的更新方向。预训练实验表明更激进的学习率衰减可在训练初期加速MuonH但可能损害后期性能。因此,保持恒定角速度并不能消除学习率调度问题,谨慎调度对实现超球风格优化器的潜力至关重要。

英文摘要

For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing updates. However, the source of their advantage remains unclear. Starting from the angular displacement between consecutive parameter states, we derive an angular effective learning rate that accounts for the parameter-update angle, parameter norm, and update norm. We also show that the conventional norm-based measure is a special case under parameter-update orthogonality. We then decompose optimizer updates into radial and tangential components and analyze how radial updates affect one-step angular displacement. Under the training configurations considered, numerical results show that the radial component has only a limited direct effect on the angular effective learning rate. It therefore cannot explain why MuonH converges more slowly than MuonWD early in training but overtakes it later. To further isolate the underlying mechanism, we devise a heuristic experiment that modifies only the learning-rate schedule so that the dynamics of each optimizer reproduce those of the other. The results suggest that their main difference stems from the evolution of the effective step size rather than an intrinsically superior update direction induced by Hyperball. Our pretraining experiments further show that more aggressive learning-rate decay can accelerate MuonH early in training but may impair its later performance. Thus, maintaining a constant angular velocity does not eliminate the learning-rate-scheduling problem; careful scheduling remains essential to realizing the potential of Hyperball-style optimizers. Our code is publicly available at https://github.com/mangocrazz/hyperball-may-not-be-a-free-lunch.

Comments14 pages, 4 figures. Code: https://github.com/mangocrazz/hyperball-may-not-be-a-free-lunch. Equal contribution: Yihao Xiao and Jialong Sun. Corresponding author: Bryan Dai

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑