用于随机可达-规避分析的物理信息强化学习
Physics-informed Reinforcement Learning for Stochastic Reach-Avoid Analysis
浏览论文内容
中文总结 AI 辅助
本文提出物理信息强化学习(PIRL)框架,结合PINNs与RL优势,开发调度式PIRL算法,缓解传统PINN失效模式,通过两案例研究验证其用于随机可达-规避分析的有效性。
中文摘要 AI 辅助
受控动力系统的随机可达-规避分析是不确定性下安全关键控制的重要工具,其中可达-规避概率由哈密顿-雅可比偏微分方程(PDE)表征。然而,随着系统维度增加,使用传统数值方法求解该PDE会变得计算上难以处理。物理信息神经网络(PINNs)在主要通过PDE残差最小化进行训练时,可能收敛到不准确的局部极小值。强化学习(RL)提供了一种可扩展的替代方案,但其学习到的值函数可能不准确或与控制PDE不一致。本文提出一种物理信息强化学习(PIRL)框架,结合PINNs和RL的互补优势用于随机可达-规避分析。我们开发了一种调度式PIRL算法,其中时序差演员-评论家学习首先引导评论家获得可达-规避值函数的有意义近似,随后逐步引入PDE残差和边界条件损失,以确保与控制PDE及其边界条件的一致性。所提方法缓解了传统PINN技术的失效模式,同时达到与成功训练的PINNs相当的精度。通过两个案例研究验证了该框架的有效性。
英文摘要
Stochastic reach-avoid analysis of controlled dynamical systems is an important tool for safety-critical control under uncertainty, in which the reach-avoid probability is characterized by a Hamilton-Jacobi partial differential equation (PDE). However, solving this PDE using conventional numerical methods becomes computationally intractable as the system dimension increases. Physics-informed neural networks (PINNs) may converge to inaccurate local minima when trained primarily through PDE-residual minimization. Reinforcement learning (RL) offers a scalable alternative, but its learned value functions may be inaccurate or inconsistent with the governing PDE. This paper proposes a physics-informed RL (PIRL) framework that combines the complementary strengths of PINNs and RL for stochastic reach-avoid analysis. We develop a scheduled PIRL algorithm in which temporal-difference actor-critic learning first guides the critic toward a meaningful approximation of the reach-avoid value function. PDE-residual and boundary-condition losses are then introduced progressively to enforce consistency with the governing PDE and its boundary conditions. The proposed method mitigates the failure modes of conventional PINN techniques while achieving accuracy comparable to that of successfully trained PINNs. The effectiveness of the proposed framework is demonstrated through two case studies.