arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向下层非凸问题的多目标双层优化的高效无海森方法

Efficient Hessian-Free Methods for Multi-Objective Bilevel Optimization with Nonconvex Lower Level

Yicong Jiang, Feihu Huang

arXiv 2608.12704首次发表:更新:

发表机构

College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对下层非凸的多目标双层优化问题,提出MOMEHA及MB-MOMEHA算法,通过Moreau包络转化问题并结合平滑加权Tchebycheff标量化,在少样本元学习等任务上优于现有方法。

AI 中文摘要

多目标双层优化在自动化学习、多任务元学习等AI领域有广泛应用。尽管近期已有部分工作开始研究多目标双层优化,但所提方法均依赖于(强)凸的下层问题。然而,这类多目标双层学习问题通常是非凸的,尤其是其下层问题为非凸。为填补这一空白,本文提出一类基于多目标Moreau包络的无海森算法(MOMEHA),用于求解下层非凸的多目标双层学习问题。具体而言,该方法利用Moreau包络将原问题转化为带有包络约束的多目标单层优化问题;通过结合平滑加权Tchebycheff标量化,在多目标场景下保留了单循环、无海森的计算优势。此外,本文还提出MOMEHA的动量变体(即MB-MOMEHA),用于求解随机多目标双层学习问题。理论上,本文在确定性和随机两种场景下均给出了所提算法的收敛性性质。在少样本元学习和神经架构搜索上的部分实验表明,本文方法在帕累托前沿上优于现有方法,验证了其有效性和鲁棒性。

英文摘要

Multi-objective bilevel optimization has wide applications in the AI area such as automated learning and multi-task meta-learning. Although recently some works have been begun to study the multi-objective bilevel optimization, the proposed methods rely on the (strongly) convex lower level problems. In fact, these multi-objective bilevel learning problems are generally nonconvex, and particularly their lower level problems are nonconvex. To fill this gap, we propose a class of Multi-Objective Moreau Envelope based Hessian-free Algorithms (MOMEHA) for the multi-objective bilevel learning problems with nonconvex lower level. Specifically, our method uses the Moreau envelope to relax the original problem into a multi-objective single-level optimization with an envelope constraint. In particular, our method retains computational advantages of being single-loop and Hessian-free in the multi-objective setting by incorporating a smooth weighted Tchebycheff scalarization. Furthermore, we propose a momentum-based variant of MOMEHA (i.e., MB-MOMEHA) method for the stochastic multi-objective bilevel learning problems. In theory, we provide the convergence properties of our algorithms under both deterministic and stochastic setting. Some experiments on few-shot meta-learning and neural architecture search demonstrate that our methods outperform the existing approaches in Pareto front, validating its effectiveness and robustness.

Comments49 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑