arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14492cs.LG

基于低秩适配器的近似Muon优化器

Approximate Muon with low-rank adapters

  • University of Bristol(布里斯托大学)
  • School of Mathematics(数学学院)
  • School of Engineering Mathematics and Technology(工程数学与技术学院)
  • University of Bath(巴斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Ben Anson, Conor Houghton, Edward Milsom

AI总结:

本文针对Muon优化器难以与LoRA结合用于低秩微调的问题,提出sMuon方法,经实验验证其在SFT和ReLoRA预训练中表现良好,可适度提升低秩微调性能。

AI中文摘要:

Muon优化器在神经网络预训练中相比其他方法展现出明显优势,但在参数高效微调(PEFT)中使用频率较低,一个潜在原因是最常见的PEFT方法LoRA无法与Muon自然结合,因为基于低秩参数化的权重更新无法实现正交化,这在数学上不可行。本文通过线性化和最小二乘法,在低秩场景下近似求解松弛的Muon目标的解,提供仅使用矩阵乘法、无需复杂线性代数分解的高效实现。本文提出的方法sMuon(小型Muon)在监督微调(SFT)和ReLoRA预训练实验中表现良好,尽管结果依赖于模型和评估方式,但整体来看,在低秩微调中使用Muon可带来适度的性能提升。

英文摘要:

The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. However, it is used less frequently for parameter-efficient fine-tuning (PEFT). One potential reason is that the most common PEFT method, LoRA, does not naturally combine with Muon since it is not mathematically possible to orthogonalize the weight update given by a low-rank parameterization. In this paper, we address this issue by approximating the solution to a relaxed Muon objective in the low-rank setting via linearization and then least-squares. We provide an efficient implementation that uses matmul operations only, as opposed to more complex linear algebra decomposition routines. Our method, sMuon (small Muon), performs favourably across SFT and a ReLoRA pretraining experiment. While results are model- and eval-dependent, we find overall that using Muon for low-rank fine-tuning provides moderate performance improvements.

↑