arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16169cs.LGcs.AI

μ子何时有助于智能体强化学习?

When Does Muon Help Agentic Reinforcement Learning?

Kai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun

首次发表
浏览论文内容

中文总结 AI 辅助

研究μ子在稀疏奖励智能体强化学习中的作用,通过与AdamW对比,发现在组内组策略优化下,μ子应用于隐藏权重矩阵能提升成功率,效果与优势估计器、学习率有关,表明其可使智能体RL受益,还需多种子和跨任务验证。

中文摘要 AI 辅助

μ子在大规模预训练中与AdamW具有竞争力,但其在强化学习(RL)后训练中的价值仍不明确。我们通过在ALFWorld上使用Qwen2.5 - 0.5B - Instruct与AdamW进行匹配单种子比较,研究了稀疏奖励智能体RL中的普通μ子。在组内组策略优化(GiGPO)下,仅将μ子应用于隐藏权重矩阵可使最终窗口验证成功率从0.290提高到0.546(提高88%);高速率AdamW控制则没有更新后的成功率。效果取决于优势估计器和学习率。在3e - 5时,μ子将GRPO从0.161提高到0.268,而GraphGPO的后期窗口差距在接近饱和时缩小。在1e - 5时,GraphGPO μ子达到0.901,将归一化验证AUC从0.399提高到0.556,并分别提前30和60次更新达到0.5和0.75的成功率。这些探索性结果表明,μ子可以使智能体RL受益,并促使联合研究策略优化器、优势估计器和学习率。多种子和跨任务验证仍未解决。

英文摘要

Muon is competitive with AdamW in large-scale pre-training, but its operating regime in reinforcement-learning post-training remains unclear. We map this regime on ALFWorld, a sparse-reward agentic benchmark, using three group-based objectives and Qwen2.5 models from 0.5B to 3B. Under a shared KL and clipping recipe, matched optimizer comparisons and AdamW rate controls trace the usable step-size range. AdamW responds non-monotonically to rate, whereas fan-in Muon remains stable at a more aggressive effective step: at $3 \times 10^{-5}$ it improves late success over an AdamW $10^{-6}$ baseline after correction across rate-metric tests. Its normalized-AUC effect is directionally positive but less uniform; the heuristic-matched lower-rate effect is less consistent, and tuned AdamW nearly matches high-rate Muon at 3B GraphGPO. High-rate Muon applies $3.53 \times$ AdamW's hidden-matrix update RMS; a full-budget RMS-matched control removes the late-success gain. Together, these results identify a recipe-level operating regime in which fan-in Muon supports a more aggressive stable effective step under shared KL and clipping: the margin is largest when optimization headroom remains and contracts near saturation, after AdamW tuning, or under magnitude matching. The scale-matched control ties this spectral effect to Muon's scale convention rather than establishing a universal optimizer ranking. Code is available at https://github.com/x66ccff/verl-muon.

↑