发表机构
MIT CSAIL; New York University(麻省理工学院计算机科学与人工智能实验室; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究分析了优化器性能随训练时长的变化,对比了Muon、SOAP、ADANA与AdamW在不同参数规模、过训练因子下的表现,确立训练时长为优化器评估与设计的关键轴。
AI 中文摘要
我们研究优化器如何沿过训练轴扩展,结果表明,优化器的相对性能和最佳超参数会随训练时长发生显著变化。具体而言,我们研究矩阵预条件方法(Muon和SOAP)以及动量调度方法(ADANA)相对于AdamW的扩展情况。我们在参数规模为51M至253M的模型上,以及过训练(OT)因子为1x至256x的范围内对比这四种优化器,同时在每个设置下扫掠基础学习率。研究发现,首选的学习率调度会沿过训练轴发生反转,最佳权重衰减系数大致按OT的平方根缩放,且更长的训练时长通常更适合更长的固定内存。在针对每个时长单独调整AdamW的固定内存后,ADANA相对于AdamW的扩展优势依然存在。对数时间权重衰减和动量冷却为ADANA带来了显著增益,且这些增益会随训练时长增加而累积。采用该处理方式后,ADANA超越了AdamW,其指数优势接近幂律随机特征的DANA理论所预测的结果。Muon和SOAP则在大部分测量范围内相对于AdamW提供了大致恒定的令牌效率优势,尽管SOAP可能在最高过训练因子下获得进一步增益。ADANA起初落后于两种矩阵预条件优化器,但随着训练时长增加缩小了差距,在最高OT因子下超过Muon并与SOAP具有竞争力。这些结果确立了训练时长是优化器评估和设计的一个重要轴。
英文摘要
We investigate how optimizers scale across the overtraining axis and show that relative optimizer performance and optimal hyperparameters change substantially with training horizon. In particular, we study how matrix-preconditioned methods (Muon and SOAP) and a momentum-scheduled method (ADANA) scale relative to AdamW. We compare these four optimizers across models from 51M to 253M parameters and overtraining (OT) factors from 1x to 256x, sweeping the base learning rate at every setting. The preferred learning rate schedule can reverse across the overtraining axis, the best weight decay coefficient scales approximately as sqrt(OT), and longer horizons generally favor longer fixed memory. ADANA's scaling advantage over AdamW persists after tuning AdamW's fixed memory separately at each horizon. Log-time weight decay and momentum cooldown provide substantial gains for ADANA that compound as training increases. With this treatment, ADANA outscales AdamW with an exponent advantage close to that predicted by DANA theory on power-law random features. Muon and SOAP instead provide roughly constant token-efficiency advantages over AdamW across most of the measured range, although SOAP may gain further at the highest overtraining factors. ADANA begins behind both matrix-preconditioned optimizers but closes the gaps as training increases, surpassing Muon and becoming competitive with SOAP at our highest OT factors. These results establish training horizon as an essential axis for optimizer evaluation and design.