发表机构
University of Bucharest; Bitdefender(布加勒斯特大学; 比特梵德)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明使用 Muon 优化器训练简单的增量式 LoRA 即可达到或超越专用持续学习方法,通过正交化更新分布,无需任务感知约束,揭示了优化器几何在缓解遗忘中的关键作用。
AI 中文摘要
使用低秩适配器(Low-Rank Adapters, LoRA)的持续学习通常通过惩罚新更新与累积过去权重之间的重叠来缓解遗忘,这会阻止某些更新方向,而不控制更新如何在剩余方向上分配其能量。我们询问这种限制是否必须依赖于任务,还是优化器提供的通用限制就足够了。我们使用 Muon 训练一个简单的增量式 LoRA(IncLoRA),该优化器对每次更新进行正交化,并将其与 O-LoRA 和 ELLA 在标准持续学习基准(Standard CL Benchmark)上的五个随机种子和三个任务顺序以及 TRACE 上的三个随机种子进行比较。IncLoRA+Muon 在 Standard CL 上达到了专用方法的准确率范围,并在 TRACE 上优于所有 AdamW 配置。一个更新约束机制就足够了,无论它来自损失函数还是优化器;在 Standard CL 上,第二个机制没有帮助,对于最受限制的方法,它损失了 8.4 个百分点的准确率和适应每个任务的可塑性。区分这两种优化器的不是更新的大小(在 Muon 下,更新大小是 AdamW 下的 0.91 到 2.06 倍),而是更新的分布方式。AdamW 将更新限制在 1.4 到 1.8 个有效奇异方向之间,而 Muon 将其分布在 7.0 个方向上,并且在任何跟踪的运行中两者都不重叠。因此,通常归因于专用持续学习方法的部分优势可能可以通过优化器更新的几何形状来解释。
英文摘要
Continual learning with Low-Rank Adapters (LoRA) typically mitigates forgetting by penalizing the overlap between a new update and the accumulated past weights, which discourages certain update directions without controlling how an update distributes its energy over the ones that remain. We ask whether that restriction has to be task-aware, or whether a generic one supplied by the optimizer is enough. We train a plain incremental LoRA (IncLoRA) with Muon, which orthogonalizes each update, and compare it against O-LoRA and ELLA over five seeds and three task orders on the Standard CL Benchmark and three seeds on TRACE. IncLoRA+Muon reaches the accuracy band of the dedicated methods on Standard CL and improves on every AdamW configuration on TRACE. One update-constraining mechanism is enough, whether it comes from the loss or from the optimizer; on Standard CL a second one does not help, and for the most restrictive method it costs 8.4 points of accuracy and the plasticity to fit each task. What separates the two optimizers is not the size of the update, which under Muon is 0.91 to 2.06 times that under AdamW, but how it is distributed. AdamW confines it to between 1.4 and 1.8 effective singular directions, Muon spreads it over 7.0, and the two do not overlap in any tracked run. Part of the advantage usually attributed to dedicated CL methods may therefore be explained by the geometry of the optimizer's updates.
Comments10 pages, 2 figures, 6 tables. Presented at the 5th Conference on Lifelong Learning Agents (CoLLAs), Work-in-Progress Track, 2026. Sebastian George Sincari and Bogdan Alexandru Gheorghe contributed equally