MOON:面向多任务学习的多目标正交归一化更新方法
MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning
AI总结:
针对现有多任务学习方法忽略现代架构矩阵结构、优化效率受限的问题,提出MOON方法,在谱-核范数几何下进行梯度操作,理论与实验均表明其可提升优化效率和多任务性能。
AI中文摘要:
多目标优化(MOO)已通过梯度操作缓解任务冲突,在多任务学习中取得显著成功。但现有多数方法将模型参数展平为向量,在欧氏几何下进行梯度操作,忽略了Transformer等现代架构中普遍存在的矩阵结构。本文表明,欧氏空间中的梯度操作在矩阵几何下通常无法得到最速下降方向,可能限制优化效率。基于矩阵值参数的最速下降理论,我们提出MOON(多目标正交归一化更新),在谱-核范数几何下执行梯度操作,并使用正交归一化的操作梯度进行参数更新。理论上,对于光滑非凸目标,我们证明在确定性设置下平均帕累托平稳性测度的收敛率为O(T^{-1/2}),在随机梯度下为O(T^{-1/4})。在各类基准上的实验结果显示,MOON始终能同时提升优化效率和最终多任务性能。我们的代码可在该URL获取。
英文摘要:
Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation. However, most existing methods flatten model parameters into vectors and perform gradient manipulation under Euclidean geometry, thereby overlooking the matrix structure prevalent in modern architectures such as Transformers. In this paper, we show that gradient manipulation in Euclidean space does not generally yield the steepest descent direction under matrix geometry, potentially limiting optimization efficiency. Drawing from the theory of steepest descent for matrix-valued parameters, we propose MOON (Multi-Objective OrthoNormalized Updates), which performs gradient manipulation under spectral--nuclear norm geometry and uses the orthonormalized manipulated gradient for parameter updates. Theoretically, for smooth non-convex objectives, we establish convergence of the averaged Pareto-stationarity measure at rates of $\mathcal{O}(T^{-1/2})$ in the deterministic setting and $\mathcal{O}(T^{-1/4})$ under stochastic gradients. Empirical results across various benchmarks show that MOON consistently improves both optimization efficiency and final multi-task performance. Our code is available at https://github.com/KunlinLyu/MOON.