AI 中文总结
针对LLM预训练中梯度噪声与病态损失景观问题,提出带球面约束的曲率条件多尺度动量方法,可加速Muon在不同架构和规模模型上的训练,兼具理论验证与实验支撑。
AI 中文摘要
预训练占据了大语言模型(LLM)训练总计算成本的很大一部分。然而,以噪声为主的梯度和高度病态的损失景观带来了严峻挑战。尽管AdamW和Muon等现代自适应优化器在大规模预训练中取得了巨大成功,但它们对梯度归一化的依赖仅能有限缓解病态曲率。沿平坦方向(小特征值的特征方向)的进展主导了最终的损失降低,但进展相对缓慢。为增强沿平坦方向的训练动态,我们提出了一种带球面约束的曲率条件多尺度动量方法,可在LLM预训练中实现稳定加速。该多尺度动量仅应用于平坦方向,将用于降噪的慢衰减分量与用于快速曲率适应的快衰减分量相结合,利用二者的互补优势。关键在于,我们采用球面约束技术防止参数膨胀和过快的有效学习率衰减,而这是简单组合会导致的问题。大量实验表明,所提方法在多种架构(密集型、混合专家MoE)和模型规模(0.12B至23亿参数)下均显著加速了Muon的训练。理论上,我们验证了加速效果,并为平坦方向多尺度动量的设计原理提供了见解。
英文摘要
Pretraining accounts for a large fraction of the total computational cost in LLM training. However, noise-dominant gradients and the highly ill-conditioned loss landscape bring severe challenges. Although modern adaptive optimizers such as AdamW and Muon have achieved great success in large-scale pretraining, their reliance on gradient normalization offers limited mitigation of the ill-conditioned curvature. The progress along flat directions (eigen-directions of small eigenvalues), which dominates the final loss reduction, remains relatively slow. To enhance training dynamics along flat directions, we propose a curvature-conditioned multiscale momentum method with sphere constraints, delivering steady acceleration in LLM pretraining. This multiscale momentum, applied only along flat directions, pairs a slow-decay component for noise reduction with a fast-decay component for rapid curvature adaptation, harnessing their complementary strengths. Crucially, we employ a sphere constraint technique to prevent parameter inflation and excessively rapid effective learning rate decay that would otherwise arise from a naive combination. Extensive experiments show that the proposed method significantly accelerates Muon across diverse architectures (dense, MoE) and model sizes (0.12B--2.3B parameters). Theoretically, we verify the acceleration effect and provide insight into the design principles underlying the flat-direction multiscale momentum.
Comments50 pages