arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

归一化-再预处理:面向LLM训练的边际尺度与交互几何的分层方法

Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training

Zixuan Gong, Zeyu Gan, Jiaye Teng, Yong Liu

arXiv 2609.36692首次发表:更新:

发表机构

Gaoling School of Artificial Intelligence; Renmin University of China; School of Statistics and Management; Shanghai University of Finance and Economics(高瓴人工智能学院; 中国人民大学; 统计与管理学院; 上海财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Normalize-Then-Precondition分层框架,通过先对角归一化再谱预处理交互几何,开发NormPre优化器,在GPT-2、LLaMA和Qwen3预训练中超越AdamW、Muon和MANO,并提供收敛保证。

AI 中文摘要

矩阵优化器已成为一个有前景的研究方向,其中Muon作为一种突出的设计脱颖而出。通过其全Gram表示重新审视Muon,我们观察到它联合处理了边际尺度和交互信息。这开辟了一种分层组织几何信息的替代方式,从而激发了归一化-再预处理(Normalize-Then-Precondition)框架。具体而言,该框架首先利用对角Gram信息构造边际归一化更新,然后对其方向性交互几何应用谱预处理。基于此框架,我们开发了NormPre,其中NormPre-G和NormPre-L分别采用全局和局部谱预处理,分别基于谱范数最速下降和正则化公式后接主导模式选择。为支持大规模训练,NormPre-G使用Newton-Schulz迭代,NormPre-L采用随机草图法来近似主导交互特征空间。理论上,我们为NormPre的简化版本建立了$\mathcal{O}(T^{-1/2})$收敛保证。在GPT-2 Small、LLaMA和Qwen3上的广泛预训练实验中,两种变体在匹配的训练预算下均持续优于AdamW、Muon和MANO。进一步的效率和谱分析揭示了两种变体的互补优势,并刻画了它们的性能-效率权衡。我们通过GitHub仓库开源了代码,网址为https://this https URL。

英文摘要

Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram representation, we observe that it jointly processes marginal-scale and interaction information. This opens an alternative way to organize geometric information hierarchically, motivating the Normalize-Then-Precondition framework. Specifically, it first uses diagonal-Gram information to construct a marginally normalized update, then applies spectral preconditioning to its directional interaction geometry. Building on this framework, we develop NormPre with NormPre-G and NormPre-L adopting global and localized spectral preconditioning, grounded in spectral-norm steepest descent and a regularized formulation followed by leading mode selection, respectively. To enable large-scale training, NormPre-G uses Newton-Schulz iterations and NormPre-L employs randomized sketching to approximate the leading interaction eigenspace. Theoretically, we establish $\mathcal{O}(T^{-1/2})$ convergence guarantees for simplified versions of NormPre. Across extensive pretraining experiments on GPT-2 Small, LLaMA and Qwen3, both variants consistently outperform AdamW, Muon and MANO under matched training budgets. Further efficiency and spectral analyses reveal the complementary strengths of two variants and characterize their performance-efficiency trade-off. We open-source our code through a GitHub repository at https://github.com/zx-gong/NormPre.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑