弱凸性与近端束方法用于鲁棒控制中的非光滑策略优化
Weak Convexity and Proximal Bundle Methods for Nonsmooth Policy Optimization in Robust Control
浏览论文内容
中文总结 AI 辅助
针对鲁棒$\mathcal{H}_\infty$控制中的非光滑策略优化,提出首个可行性保持的近端束算法,利用弱凸性和弱Polyak--Łojasiewicz不等式,实现确定性非渐近复杂度保证。
中文摘要 AI 辅助
我们研究了具有静态输出反馈的离散时间鲁棒$\mathcal{H}_\infty$控制的策略优化,并提出了第一个具有确定性、非渐近复杂度保证的可行性保持算法。该问题自然导致在稳定反馈增益集合上的非光滑和非凸优化。我们首先建立了$\mathcal{H}_\infty$代价的几个结构性质。特别地,我们证明了该代价在子水平集的每个凸子集上是弱凸的。对于状态反馈情形,我们进一步建立了弱Polyak--Łojasiewicz不等式,该不等式确保每个驻点都是全局最优的。基于这些性质,我们开发了一种用于$\mathcal{H}_\infty$策略优化的近端束方法。所提出的方法可以被视为近端点方法的可实现近似,并且仅使用函数值和次梯度信息。我们证明了所有迭代都保持稳定性,并建立了寻找$(\eta,\epsilon)$-驻点的确定性非渐近复杂度界$\mathcal{O}(\max\{\eta^{-4},\epsilon^{-2}\})$。数值实验验证了我们的理论结果。
英文摘要
We study policy optimization for discrete-time robust $\mathcal{H}_\infty$ control with static output-feedback, and present the first feasibility-preserving algorithm with a deterministic, non-asymptotic complexity guarantee. This problem naturally leads to a nonsmooth and nonconvex optimization over the set of stabilizing feedback gains. We first establish several structural properties of the $\mathcal{H}_\infty$ cost. In particular, we show that the cost is weakly convex on every convex subset of a sublevel set. For the state-feedback case, we further establish a weak Polyak--Łojasiewicz inequality, which ensures that every stationary point is globally optimal. Building on these properties, we develop a proximal bundle method for $\mathcal{H}_\infty$ policy optimization. The proposed method can be viewed as an implementable approximation of the proximal point method and uses only function value and subgradient information. We show that all iterates remain stabilizing and establish a deterministic non-asymptotic complexity bound of $\mathcal{O}(\max\{η^{-4},ε^{-2}\})$ for finding an $(η,ε)$-stationary point. Numerical experiments illustrate our theoretical results.
发表机构
- University of California San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。