arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ORCA:用于更快、更好LLM训练的退火谱条件优化器

ORCA: The Annealed Spectral Conditioning Optimizer for Faster, Better LLM Training

Yuanshi Liu, Boyuan Jiang, Liang Hou, Xin Tao, Pengfei Wan, Zhouchen Lin, Cong Fang

arXiv 2610.06116首次发表:更新:

发表机构

Peking University; Kling Team(北京大学; Kling 团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ORCA通过在训练早期施加临时软正交正则化后移除,使权重先受益于宽谱再自由适应,在多种LLM上比Muon取得更低验证损失,且无需架构改动、开销极小。

AI 中文摘要

现代LLM优化器(如Muon)通常比Adam产生更高有效秩的权重矩阵,然而进一步的谱控制仅带来适度提升。我们识别出这一结果背后的张力:集中的谱可以抑制耦合权重矩阵中的梯度方向并减慢优化,而整个训练过程中维持的约束可能限制任务特定适应性并提高可达到的损失下限。我们引入ORCA(正交正则化,冷却后),一种最小优化器干预,在训练早期应用强但临时的软正交正则化,然后移除它。这使得权重在早期受益于更广泛的谱,并在之后自由适应。在从130M到8B参数的LLaMA、Qwen3和细粒度混合专家模型中,ORCA实现了比Muon更低的最终验证损失。其相对于Muon的损失减少等于或超过Muon相对于Adam的减少。消融研究支持早期塑造、后期释放的设计。此外,ORCA不需要架构更改,且增加的开销极小。

英文摘要

Modern LLM optimizers such as Muon often produce weight matrices with higher effective rank than Adam, yet further spectral control has delivered only modest gains. We identify a tension behind this result: concentrated spectra can suppress gradient directions in coupled weight matrices and slow optimization, while constraints maintained throughout training can limit task-specific adaptation and raise the attainable loss floor. We introduce ORCA (Orthogonal Regularization, Cooled After), a minimal optimizer intervention that applies strong but temporary soft orthogonality regularization early in training, then removes it. This allows the weights to benefit from a broader spectrum early on and adapt freely afterward. Across LLaMA, Qwen3, and fine-grained mixture-of-experts models ranging from 130M to 8B parameters, ORCA achieves lower final validation loss than Muon. Its loss reduction relative to Muon matches or exceeds Muon's reduction relative to Adam. Ablations support the early-shaping, later-release design. Further, ORCA requires no architectural changes and adds minimal overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑