并行策略梯度方法用于非线性反馈控制器参数优化
Parallel Policy-Gradient Methods for Parameter Optimization of Nonlinear Feedback Controllers
- University of New Mexico(新墨西哥大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出一种时间并行策略梯度框架,用于非线性反馈控制器的参数优化,通过残差最小化和高斯-牛顿迭代加速梯度评估,并在惯性轮摆示例中验证了性能提升与计算优势。
中文摘要 AI 辅助
结构化反馈控制器提供严格的稳定性保证,但通常需要手动调整参数以实现良好的闭环性能。策略梯度方法为参数优化提供了一种系统化途径;然而,传统的梯度评估需要顺序的前向状态滚动和反向协态传播。本文针对离散时间非线性控制仿射系统,提出了一种时间并行策略梯度框架。我们推导了策略梯度表达式,其中策略梯度评估所需的状态和协态滚动被表述为残差最小化问题,并通过带有并行关联扫描的高斯-牛顿(GN)迭代求解。对于全局渐近稳定且局部指数稳定的闭环系统,我们证明了残差最小化问题满足局部Polyak-Lojasiewicz(PL)不等式,且GN迭代在局部以二次速率收敛。此外,PL常数、收敛邻域的大小以及二次收敛界均与滚动时域T无关。我们还证明,对于任何有限时域T,状态求解器从任意初始化出发,最多经过T次迭代即可恢复精确轨迹。最后,一个带有互联和阻尼分配无源控制(IDA-PBC)的惯性轮摆示例展示了所提出的并行策略梯度框架在改进闭环性能和计算优势方面的效果。
英文摘要
Structured feedback controllers provide rigorous stability guarantees, but often require manual parameter tuning to achieve good closed-loop performance. Policy-gradient methods offer a systematic approach to parameter optimization; however, conventional gradient evaluation requires sequential forward state rollout and backward costate propagation. This letter develops a time-parallel policy-gradient framework for discrete-time nonlinear control-affine systems. We derive the policy-gradient expression where the state and costate rollouts required for policy-gradient evaluation are formulated as residual-minimization problems and solved using Gauss-Newton (GN) iterations with parallel associative scans. For closed-loop systems that are globally asymptotically stable and locally exponentially stable, we show that the residual-minimization problems satisfy a local Polyak-Lojasiewicz (PL) inequality and that the GN iterates converge locally at a quadratic rate. Moreover, the PL constant, the size of the convergence neighborhood, and the quadratic convergence bound are independent of the rollout horizon T. We also prove that, for any finite horizon T, the state solver recovers the exact trajectory from any initialization in at most T iterations. Finally, an inertia-wheel pendulum example with interconnection and damping assignment passivity-based control (IDA-PBC) demonstrates improved closed-loop performance and the computational benefits of the proposed parallel policy-gradient framework.