AI 中文总结
针对非光滑 $H_\infty$ 输出反馈策略搜索,证明在扩展凸提升的精确切片上 $\varepsilon$-平稳性导致 $O(\varepsilon)$-次优性,给出收敛速率保证,并提出二分法构造显式 $\varepsilon$-最优镇定控制器。
AI 中文摘要
我们研究连续时间全阶动态输出反馈 $H_\infty$ 策略搜索,这是一个非凸且非光滑的问题。直接策略搜索是强化学习和连续控制中的核心范式,但在鲁棒输出反馈设置中,严格的理论保证仍然稀缺。$H_\infty$ 问题是一个典型的基准问题,因为它既捕捉了扰动衰减和鲁棒性,又暴露了策略空间优化的困难非光滑几何结构。我们证明,在扩展凸提升的精确恒等规范切片上,$\varepsilon$-平稳性在紧致精确切片上产生 $O(\varepsilon)$-次优性,这进而为非光滑策略搜索方法提供了收敛速率保证。该结果解决了 Guo 和 Hu [2022] 在更一般的动态输出反馈 $H_\infty$ 策略搜索设置中提出的有限时间最优性差距问题。我们进一步利用扩展凸提升所提供的价值等价性,提出了一种非严格可行性二分法,并附带一个最终的严格可行性恢复步骤,从而得到一个显式的 $\varepsilon$-最优镇定控制器。这些结果为非光滑 $H_\infty$ 策略搜索的先前定性最优性理论提供了定量和算法上的加强。
英文摘要
We study continuous-time full-order dynamic output-feedback $H_\infty$ policy search, a nonconvex and nonsmooth problem. Direct policy search is a central paradigm in reinforcement learning and continuous control, but rigorous guarantees remain scarce in robust output-feedback settings. The $H_\infty$ problem is a canonical benchmark because it captures disturbance attenuation and robustness while exposing the hard nonsmooth geometry of policy-space optimization. We prove that on the exact identity-gauge slice of the extended convex lift, $\varepsilon$-stationarity yields $O(\varepsilon)$-suboptimality on compact exact slices, which in turn yields convergence-rate guarantees for nonsmooth policy-search methods. This result addresses the finite-time optimality-gap question raised by Guo and Hu [2022] in the more general dynamic output-feedback $H_\infty$ policy-search setting. We further use the established value equivalence supplied by extended convex lifting to formulate a nonstrict-feasibility bisection method with one final strict-feasibility recovery step, yielding an explicit $\varepsilon$-optimal stabilizing controller. These results provide a quantitative and algorithmic strengthening of prior qualitative optimality theory for nonsmooth $H_\infty$ policy search.
CommentsAppeared at The 65th IEEE Conference on Decision and Control (CDC), 2026