发表机构
University of Rochester(罗切斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
多目标学习中,随机MGDA方法因小批量采样引入噪声而收敛率次优。本文指出CA方向关于雅可比矩阵是1/2-Hölder连续的,在额外正则性条件下可改进为Lipschitz连续。基于此提出随机多目标正则性感知方法,提升收敛率并建立冲突避免保证,实验验证了其有效性。
AI 中文摘要
多目标学习旨在同时优化多个目标。多梯度下降算法(MGDA)是通过沿公共下降或冲突避免(CA)方向迭代更新来实现的。然而,在随机设置中,普通随机MGDA方法(SMG)由于小批量采样在梯度中引入噪声而缺乏快速收敛率,这导致更新方向出现偏差。本文表明CA方向关于雅可比矩阵是1/2-Hölder连续的,且在最坏情况下指数1/2无法改进,这导致先前工作中普通随机MGDA的收敛率次优。在额外正则性条件下可改进为Lipschitz连续。基于此,提出了一种随机多目标正则性感知(MoRe)方法,当子问题正则时利用CA方向的Lipschitz连续性,否则切换到固定标量化权重。理论上,该方法在非凸设置中将SMG的收敛率从\(\widetilde{\mathcal O}(T^{-1/4})\)提高到\(\widetilde{\mathcal O}(T^{-1/2})\)并建立了每次迭代的冲突避免保证。实验证明了其在多任务性能方面的有效性并验证了与理论速率一致的收敛行为。
英文摘要
Multi-objective learning (MOL) aims to optimize multiple objectives simultaneously. The multi-gradient descent algorithm (MGDA) is a workhorse that iteratively updates along a common descent or conflict-avoidant (CA) direction across objectives. In stochastic settings, however, the vanilla stochastic MGDA method, SMG, lacks a fast convergence rate because mini-batch sampling introduces noise in the gradients. This causes bias in the update direction, which is controlled by the CA direction continuity. In this paper, we show that the CA direction is $1/2$-Holder continuous with respect to the Jacobian matrix, and the exponent $1/2$ cannot be improved in the worst case. This leads to a suboptimal convergence rate for vanilla stochastic MGDA in prior works. Nevertheless, under additional regularity conditions, we show this can be improved to Lipschitz continuity. Based on this insight, we propose a stochastic multi-objective regularity-aware (MoRe) method that exploits the Lipschitz continuity of the CA direction when the subproblem is regular, and switches to a fixed scalarization weight otherwise. Intuitively, the proposed algorithm employs CA direction update when the gradient conflict is large, and linear scalarization update otherwise. Theoretically, our method improves the convergence rate of SMG in the nonconvex setting from $\widetilde{\mathcal O}(T^{-1/4})$ to $\widetilde{\mathcal O}(T^{-1/2})$, where $\widetilde{\mathcal O}(\cdot)$ hides logarithmic factors. Meanwhile, we also establish the per-iterate conflict-avoidance guarantees. Empirically, experiments demonstrate its effectiveness in multi-task performance and verify convergence behavior consistent with the established theoretical rate.