基于凸提升的零阶非光滑非凸优化及其在状态反馈$H_\infty$策略优化中的应用
Zeroth-Order Nonsmooth Nonconvex Optimization with Convex Liftings and Its Application to State-Feedback $H_\infty$ Policy Optimization
浏览论文内容
中文总结 AI 辅助
本文提出一种基于凸提升的零阶近端点算法,用于求解具有凸提升结构的非光滑非凸优化问题,并将其应用于离散时间状态反馈$H_\infty$策略优化,取得了理论复杂度保证。
中文摘要 AI 辅助
直接策略优化广泛应用于强化学习与控制领域,但通常会导致非凸优化问题。对于状态反馈$H_\infty$控制,其策略目标函数是非光滑的,不过存在良性的优化景观,可通过最新提出的扩展凸提升框架揭示其隐藏凸性。受隐藏凸优化最新进展的启发,本文研究了具有凸提升结构的非光滑非凸问题的零阶优化,提出了一种零阶近端点算法:外层循环采用不精确近端点法构造强凸子问题,内层循环仅通过函数评估近似求解每个子问题。在概率至少为$1-\delta$的情况下,该算法返回$\epsilon$-最优解所需的函数评估次数为$\widetilde{O}\left(d\epsilon^{-3}\right)$,且所有迭代点均保持可行,无需显式投影。最后,本文验证了分析所需假设适用于离散时间状态反馈$H_\infty$策略优化,达到指定目标值间隙的神谕复杂度为$\widetilde{O}\left(n_u n_x\epsilon^{-3}\right)$,其中$n_u \times n_x$为待优化反馈增益的维度。
英文摘要
Direct policy optimization is widely used in reinforcement learning and control, but generally leads to nonconvex optimization problems. For state-feedback $H_\infty$ control, the policy objective is also nonsmooth, despite possessing a benign landscape whose hidden convexity can be revealed by the recently developed extended convex lifting framework. Motivated by recent advances in hidden convex optimization, we study zeroth-order optimization of nonsmooth, nonconvex problems admitting a convex lifting. We propose a zeroth-order proximal point algorithm: An inexact proximal-point outer loop constructs strongly convex subproblems, while an inner loop approximately solves each subproblem using only function evaluations. With probability at least $1-δ$, our proposed algorithm returns an $ε$-optimal solution using $\widetilde{O}\left(dε^{-3}\right)$ function evaluations, while all iterates remain feasible without explicit projection. Finally, we verify that the assumptions underlying our analysis hold for discrete-time state-feedback $H_\infty$ policy optimization, yielding an oracle complexity of $\widetilde{O}\left(n_u n_xε^{-3}\right)$ for attaining a prescribed objective value gap, where $n_u\times n_x$ is the dimension of the feedback gain to be optimized over.