arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于Muon优化器的物理响应-记忆模型

A Physical Response-and-Memory Model for Muon Optimization

Yinze Hu, Hongjun Xiang, Xingao Gong, Hongyu Yu

arXiv 2608.22994首次发表:更新:

发表机构

School of Physics Science and Engineering, Tongji University(同济大学物理科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文构建训练权重矩阵的物理响应-记忆模型,解释Muon优化器半正交化方向的有效性,提出Bi-Maxwell优化器,其双时间尺度记忆核可减少大型语言模型训练步数。

AI 中文摘要

训练大型语言模型的成本很高。相同计算量下最终能达到的损失值取决于每一步梯度如何转化为权重更新,这种转化规则就是优化器。从SGD、AdamW到近期的Muon,有效的更新规则大多由工程直觉塑造,再在基准上进行选择。Muon在应用更新前会对动量矩阵进行半正交化,且在公开训练基准上不断打破纪录;但半正交化方向为何有效,以及动量应平均多长时间的历史,目前主要靠经验回答。本文将训练过程中的权重矩阵视为具有记忆的响应介质,并为其构建物理模型,两个问题在此模型中得到了答案:半正交化方向是输出侧安全预算下的最大耗散响应,这解释了其为何有效;动量是介质积累的内部应力,其平均时长由该应力的弛豫决定,真实介质在多个时间尺度上弛豫,最简单的形式是一个快时间尺度和一个慢时间尺度。在此基础上,我们提出了Bi-Maxwell优化器。该框架进一步产生了可验证的结论:梯度方向在训练早期变化快,后期变化慢,因此最优记忆长度应随训练阶段增长;我们通过只读探针在8条独立训练轨迹上对其代理进行的逐步测量与该结论一致。仅将记忆核从单时间尺度改为双时间尺度,就能在公开大型语言模型优化器基准上,以明显更少的步数将训练达到目标损失。

英文摘要

Training large language models is costly. How low a loss the same compute can ultimately reach depends on how each step's gradient is converted into a weight update; the rule that performs this conversion is the optimizer. From SGD and AdamW to the recent Muon, effective update rules have mostly been shaped by engineering intuition and then selected on benchmarks. Muon semi-orthogonalizes the momentum matrix before applying the update and has kept breaking records on public training benchmarks; yet why the semi-orthogonalized direction works, and over how long a history the momentum should average, are two questions at present answered mainly by experience. Here we treat the weight matrix during training as a responsive medium with memory and build a physical model for it, in which both questions find answers: the semi-orthogonalized direction is the maximally dissipative response under an output-side safety budget, which explains why it works; momentum is the internal stress accumulated by the medium; how long it should average is set by the relaxation of this stress, and a real medium relaxes on more than one timescale, the simplest form being one fast and one slow. On this basis we propose the Bi-Maxwell optimizer. The framework further yields a testable consequence: gradient directions change fast early in training and more slowly later, so the optimal memory length should grow with training stage; step-by-step measurements of a proxy for it by a read-only probe across 8 independent training trajectories are consistent with this consequence. Replacing the memory kernel alone, from a single timescale to two, brings training to the target loss in noticeably fewer steps on a public large-language-model optimizer benchmark.

Comments44 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑