arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10975cs.LGcs.AI

用链式LMO优化大语言模型

Optimizing Large Language Models with Chained LMOs

Sungyoon Kim, Kaan Ozkara, Youngsuk Park

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对现有Muon类优化器缺乏统一视角的问题,提出链式LMO框架及新型优化器TensorChain,在Qwen3预训练中实现了优于基线的token效率与9.6%的token节省。

中文摘要 AI 辅助

Muon已催生了一系列由多个矩阵归一化组成的优化器,但这些方法仍零散且缺乏统一视角。我们引入链式线性最小化预言机(chained LMOs),将这些方法建模为LMO的组合。尽管这些链式方法在经验上取得成功,但许多链式结构超出了标准LMO框架,且在平滑凸目标上可能发散。为解释组合为何仍有帮助,我们转向线性关联记忆,证明在各向异性嵌入下,链式结构可优于Muon。在实验中,我们提出框架内的新型优化器TensorChain,其堆叠不同层的兼容权重矩阵,并对三维张量沿各轴进行归一化。在Qwen3 0.6B和1.7B预训练中,TensorChain在平均token效率上优于所有链式基线,在匹配验证损失下,相比Muon平均token节省9.6%。

英文摘要

Muon has motivated a growing family of optimizers that compose multiple matrix normalizations, but these methods remain fragmented and lack a unified perspective. We introduce chained linear minimization oracles (chained LMOs), which cast these methods as compositions of LMOs. Despite their empirical success, many chains fall outside the standard LMO framework and can diverge on smooth convex objectives. To explain why composition can nevertheless help, we turn to linear associative memory and show that chaining can improve over Muon under anisotropic embeddings. Empirically, we propose TensorChain, a novel optimizer within the framework that stacks compatible weight matrices across different layers and normalizes the 3d tensor across its axes. In Qwen3 0.6B and 1.7B pretraining, TensorChain outperforms all chained baselines in average token efficiency, with average token savings of 9.6% over Muon at matched validation loss.

发表机构

  • Stanford University(斯坦福大学)
  • Amazon Annapurna Labs(亚马逊安纳普尔纳实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑