发表机构
School of Mathematics and Statistics, Xi’an Jiaotong University; School of Data Science, Fudan University; SGIT AI Lab, State Grid Corporation of China; Center for Intelligent Decision-Making and Machine Learning, School of Management, Xi’an Jiaotong University(西安交通大学数学与统计学院; 复旦大学数据科学学院; 国家电网公司SGIT人工智能实验室; 西安交通大学管理学院智能决策与机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究双层优化中理论与实际单环实现的差距,利用解耦范数分析框架,改进单环近似隐式微分和迭代微分方法的收敛结果,提升收敛速率并精确匹配渐近误差下界,经实验验证理论发现。
AI 中文摘要
双层优化支撑着许多机器学习应用,如超参数优化、元学习、神经架构搜索和强化学习。基于超梯度的方法虽有显著进展,但理论保证与高效的实际单环实现之间仍存在差距。我们利用提出的解耦范数分析(DNA)框架,为单环近似隐式微分(AID)和迭代微分(ITD)方法建立了更精确的收敛结果,弥合了这一差距。对于AID,将收敛速率从\(\mathcal{O}(\kappa^6/K)\)提高到\(\mathcal{O}(\kappa^5/K)\);对于ITD,证明渐近误差为\(\mathcal{O}(\kappa^2)\),与已知下界精确匹配并改进了先前的\(\mathcal{O}(\kappa^3)\)保证。合成和实际任务的数值实验证实了我们的理论发现。
英文摘要
Bilevel optimization underpins many machine learning applications, including hyperparameter optimization, meta-learning, neural architecture search, and reinforcement learning. While hypergradient-based methods have advanced significantly, a gap persists between theoretical guarantees and practical single-loop implementations required for efficiency. We bridge this gap by establishing sharper convergence results for single-loop approximate implicit differentiation (AID) and iterative differentiation (ITD) methods, leveraging our proposed analytical framework, decoupled norm analysis (DNA). For AID, we improve the convergence rate from $\mathcal{O}(κ^6/K)$ to $\mathcal{O}(κ^5/K)$, where $κ$ is the condition number of the inner-level problem. For ITD, we prove that the asymptotic error is $\mathcal{O}(κ^2)$, exactly matching the known lower bound and improving upon the previous $\mathcal{O}(κ^3)$ guarantee. Numerical experiments on synthetic and real tasks corroborate our theoretical findings.
Comments26 pages,6 figures