发表机构
Politecnico di Milano; MOX - Department of Mathematics, Politecnico di Milano(米兰理工大学; 米兰理工大学数学MOX系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对高维和参数化动态系统控制,提出物理增强强化学习(PEARL)范式,利用其动力学可微性,采用演员-伴随算法,经实验验证该方法能减少环境交互、样本效率高、可泛化并能扩展到高维空间。
AI 中文摘要
强化学习(RL)最近成为用于非线性和复杂动态系统的一种有前景的反馈控制策略。然而,RL算法样本效率低,需要大量与环境的交互来合成最优控制策略。由于高维空间中探索-利用困境带来的维度诅咒,RL应用通常限于稀疏传感器和执行器。在这项工作中,我们用一种新颖的物理增强强化学习(PEARL)范式将RL与传统最优控制相联系,该范式针对高维和参数化动态系统控制进行了定制,利用其动力学的可微性。具体而言,PEARL采用一种演员-伴随算法,利用自动微分在短时间范围内计算策略梯度,并通过神经网络近似未来回报的基于伴随的敏感性,显著减少环境交互次数,同时减轻长期梯度不稳定性。通过两个具有挑战性的非定常流中的参数化导航问题,我们表明PEARL:(i)有效利用可微环境优于现有RL算法;(ii)由于物理引导的策略学习,样本效率高;(iii)能在多种场景中泛化,这在处理参数化系统时至关重要;(iv)能将RL扩展到高维状态和动作空间,无需低维状态表示或多智能体策略。
英文摘要
Reinforcement learning (RL) has recently emerged as a promising feedback control strategy for nonlinear and complex dynamical systems. However, RL algorithms are sample inefficient and require a large number of interaction with the environment to synthesize optimal control strategies. Consequently, applications of RL are typically limited to sparse sensors and actuators due to the curse of dimensionality entailed by the exploration-exploitation dilemma in high-dimensional spaces. In this work, we bridge RL and traditional optimal control for dynamical system with a novel Physics-EnhAnced Reinforcement Learning (PEARL) paradigm tailored to the control of high-dimensional and parametric dynamical systems, exploiting the differentibility of their dynamics. Specifically, PEARL employs an actor-adjoint algorithm that leverages automatic differentiation to compute policy gradients over short horizons and adjoint-based sensitivities of future returns approximated via neural networks, significantly reducing the number of environment interactions, while mitigating long-term gradient instabilities. Through two challenging parametric navigation problems in unsteady flows, we show that PEARL (i) effectively exploits differentiable environments to outperform state-of-the-art RL algorithms, (ii) is sample efficient, thanks to the physics-guided policy learning, (iii) generalizes across multiple scenarios, which is crucial when dealing with parametric systems, and (iv) enables scaling RL to high-dimensional state and action spaces, without requiring low-dimensional state representations or multi-agent strategies.