M+Adam:通过加法-乘法优化进行低精度训练
M+Adam: Low-Precision Training via Additive-Multiplicative Optimization
浏览论文内容
中文总结 AI 辅助
研究低精度训练中标准优化器的问题,提出结合加法和乘法更新的M+Adam方法,经证明在标准假设下单调下降,在多种模型预训练中持续改善低精度训练。
中文摘要 AI 辅助
使用量化权重训练可降低成本,但常导致精度下降,尤其在低精度优化且不存储高精度副本时。我们识别出关键失败模式:低精度下标准优化器会卡住。此前提出乘法更新取代加法更新,虽在极低精度成功,但在零附近和符号变化处失败。加法和乘法更新的失败模式互补,因此我们提出M+Adam,结合两者。在标准平滑假设下证明其单调下降。在多种模型预训练中,M+Adam持续改善低精度训练。
英文摘要
Training with quantized weights can reduce costs but often results in degraded accuracy, especially when optimization is carried out in low precision, without storing high-precision copies. We identify a key failure mode: under low precision, standard optimizers can get stuck and not make progress, especially at large weight magnitudes due to coarse mantissa resolution. To overcome this, multiplicative updates have been previously proposed, in place of additive updates in standard optimizers. While successful under extremely low precision, such as under the logarithmic number system, they suffer from failures near zero and across sign changes. The failure modes of additive and multiplicative updates are therefore complementary. To exploit this, we propose M+Adam, which combines both update types: additive steps handle sign changes and small magnitudes, while multiplicative steps ensure progress at large magnitudes when additive updates are zeroed out under rounding. We prove monotone descent for M+Adam under standard smoothness assumptions. Across LLaMA-style pretraining with 60M-1B models, 1x-8x Chinchilla budgets, and using only BF16, FP8, and FP4 master weights, M+Adam consistently improves low-precision training.
发表机构
- California Institute of Technology(加利福尼亚理工学院)
- University of Copenhagen(哥本哈根大学)
- Aarhus University(阿鲁斯大学)
机构由 AI 辅助整理,请以论文原文为准。