arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27518cs.LG

当μ子遇见任务干扰:持续学习与模型融合的谱视角

When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging

Shangge Liu, Yuehan Yin, Yinghuan Shi, Lei Wang, Wenbin Li

首次发表
浏览论文内容

中文总结 AI 辅助

本研究将持续学习与模型融合的共同挑战归为任务干扰,发现Muon优化器可通过调控谱范数收紧干扰界,在多任务基准上替换AdamW能显著提升性能。

中文摘要 AI 辅助

持续学习(CL)和模型融合(MM)均旨在获得一个能在多个任务上表现良好的单一模型,二者分别面临灾难性遗忘和权重解缠误差的挑战。现有文献中,这些困难仅被单独处理并通过多种方案缓解,而基础优化器诱导的几何结构被视为实现细节。本研究表明,这两种困难实际是同一现象的两个实例:对一个任务有用的参数更新会改变模型在另一个任务上的输出。我们将这一共性现象形式化为「任务干扰」,并将其简化为一个通用的层间Frobenius内积⟨ΔW_ℓ, J_ℓ(x)⟩_F。该量随后被用于揭示优化器的作用。我们从理论上推导了一个上界,将谱范数∥ΔW_ℓ∥₂分离为优化器可控制的任务干扰因子,且逐模式分析显示该上界可追踪经验干扰的主导部分。具体而言,我们发现近期提出的Muon优化器是一种通过构造来调控该因子的机制。本研究揭示,Muon对谱范数的精妙控制收紧了CL和MM的干扰界,使其成为一种以优化器为中心、与现有方案互补的原则性方法。我们的理论分析得到了实验结果的充分验证:在三个CLIP骨干网络的八项任务模型融合基准上,用Muon替换AdamW优化器可使准确率提升多达5.02个百分点;对于持续学习,Muon在十类增量协议、三项任务增量协议及11任务MTIL基准上均取得了一致的正向增益。

英文摘要

Continual learning (CL) and model merging (MM) both aim to obtain a single model that performs well across multiple tasks, challenged respectively by catastrophic forgetting and weight-disentanglement error. In the literature, these difficulties are merely treated separately and mitigated through a variety of solutions, while the geometry induced by the base optimizer is treated as an implementation detail. In this work, we show that the two difficulties are in fact two instances of the same phenomenon: a parameter update useful for one task shifts the model's outputs on another. We formalize this shared phenomenon as \textit{task interference} and reduce it to a common layer-wise Frobenius inner product $\langle ΔW_\ell, J_\ell(x)\rangle_F$. This quantity, in turn, is utilized to expose the role of the optimizer. We theoretically derive an upper bound that isolates the spectral norm $\|ΔW_\ell\|_2$ as an optimizer-controllable factor of task interference, and a per-mode analysis shows that this bound tracks the dominant part of the empirical interference. Specifically, we then identify the recent Muon optimizer as a mechanism that regulates this factor by construction. Our work reveals that its elegant control on spectral norm tightens the interference bound for both CL and MM, positioning Muon as a principled optimizer-centric approach complementary to existing solutions. Our theoretcal analysis is well validated by experimental results. Replacing the AdamW optimizer with Muon improves accuracy by up to +5.02 points on the eight-task model-merging benchmark across three CLIP backbones. For continual learning, Muon also delivers uniformly positive gains across ten class-incremental protocols, three task-incremental protocols, and the 11-task MTIL benchmark.

发表机构

  • State Key Laboratory for Novel Software Technology, Nanjing University(南京大学现代国家重点实验室(注:此处按原文直译,实际常见译名为“南京大学计算机软件新技术国家重点实验室”))
  • University of Wollongong(卧龙岗大学)

机构由 AI 辅助整理,请以论文原文为准。

↑