arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

最佳优化器取决于批大小

The Best Optimizer Depends on Batch Size

Xingyu Dang, Kaiyue Wen, Sadhika Malladi

arXiv 2610.08975首次发表:更新:

发表机构

Princeton University; Stanford University; University of California, San Diego(普林斯顿大学; 斯坦福大学; 加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文挑战了优化器性能与批大小无关的假设,证明Muon缺乏普适缩放规则,且语言模型预训练的最佳优化器随批大小变化而改变。

AI 中文摘要

大量新的自适应优化器被设计用于高效估计并利用小批量梯度统计信息来塑造参数更新,但它们通常在单一批大小下进行基准测试。超参数缩放规则承诺在批大小和梯度噪声变化时保持性能,这表明在某一批大小下表现最佳的优化器在另一批大小下也应保持最佳。我们通过以下两点挑战这种开发和评估优化器的方法:(1)对于Muon,没有一种有原则的缩放规则能在不同训练设置中一致地发挥作用;(2)即使在进行了广泛的超参数调优之后,语言模型预训练的最佳优化器也会随批大小的变化而改变。

英文摘要

A plethora of new adaptive optimizers are designed to efficiently estimate and use minibatch gradient statistics to shape parameter updates, but they are typically benchmarked at a single batch size. Hyperparameter scaling rules promise to preserve performance as batch size and gradient noise change, suggesting that the best optimizer at one batch size should remain the best at another. We challenge this approach to developing and evaluating optimizers by showing: (1) no principled scaling rule for Muon works consistently across training settings, and (2) the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑