arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RODE:一种用于优化的径向-正交解耦引擎

RODE: A Radial-Orthogonal Decoupled Engine for Optimization

Guoxiang Xu, Bince Qu, Qi Sun, Cheng Zhuo

arXiv 2608.21024首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出RODE优化器,通过解耦径向与方向分量的更新规则,在语言建模、图像分类等任务中优于Muon变体,可降低损失与模型范数,提升优化效果。

AI 中文摘要

现代神经网络训练越来越多地使用感知矩阵的优化器,然而它们的条件矩阵步长通常直接加到权重上,同时改变权重的范数和方向。这种相互作用很重要,因为当前的范数决定了角运动,而方向学习会驱动范数增长,从而改变后续步骤。我们引入RODE,它为径向和方向分量提供独立的更新规则和步长。RODE通过标量径向规则显式更新矩阵的Frobenius范数,而其方向通道在切空间中执行Newton-Schulz条件更新。受控的GPT-2干预实验显示,直接范数控制和RODE的方向更新均能带来增益。在两个语言建模和两个图像分类任务中,RODE在所有直接比较中均优于Muon变体,且最终全模型范数更低。在15亿参数规模下,使用从Qwen2风格语言模型搜索中直接迁移的学习率,相较于Muon RMS,RODE将损失从4.145降至3.346,最终全局范数从11964降至2183,固定半径RODE的性能进一步提升。对于Qwen3.5-9B全参数微调,所有六个优化器使用相同的调优预算、正式训练和评估设置;RODE在所有四个评估任务中均优于两种Muon变体,且在GSM8K和MATH-500上取得最高均值。因此,解耦径向和方向动态提供了一种更有效、更可控的矩阵优化方法。

英文摘要

Modern neural network training increasingly uses matrix-aware optimizers, yet their conditioned matrix step is typically added directly to the weight, jointly changing its norm and direction. This interaction matters because the current norm determines angular motion, while directional learning can drive norm growth and thereby alter later steps. We introduce RODE, which gives the radial and directional components separate update rules and step sizes. RODE explicitly updates the matrix Frobenius norm through a scalar radial rule, while its directional channel performs Newton--Schulz-conditioned updates in the tangent space. Controlled GPT-2 interventions show gains from both direct norm control and RODE's directional update. Across two language-modeling and two image-classification tasks, RODE outperforms both Muon variants in every direct comparison and ends with lower full-model norms. At 1.5B scale, using the learning rate transferred directly from the Qwen2-style LM sweep, RODE lowers loss from 4.145 to 3.346 and final global norm from 11964 to 2183 relative to Muon RMS, with fixed-radius RODE improving further. For Qwen3.5-9B full-parameter fine-tuning, all six optimizers use the same tuning budget and the same formal-training and evaluation settings; RODE outperforms both Muon variants on all four evaluation tasks and attains the highest mean on GSM8K and MATH-500. Thus, decoupling radial and directional dynamics offers a more effective and controllable approach to matrix optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑