发表机构
MOX, Department of Mathematics Politecnico di Milano; Department of Aeronautics Imperial College London(米兰理工大学数学MOX系; 伦敦帝国理工学院航空系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在测量和模型不确定下参数化动态系统的稳健控制,提出基于超网络与集成学习结合的HypEMBER框架,通过超网络表示策略和价值函数实现参数泛化,用逼近器集成量化不确定性,实验表明该框架能提升训练稳定性等,增强对不确定性的稳健性。
AI 中文摘要
在这项工作中,我们研究了强化学习(RL)作为在存在测量和模型不确定性的情况下对参数化动态系统进行稳健控制的框架。高维状态空间、昂贵的数值求解器、控制方程的部分知识以及对可能不确定或难以准确估计的物理参数的依赖,使得使用标准RL方法在计算上不可行。缺乏稳健性和跨参数变化的泛化性在存在噪声或不完整测量时会进一步放大,最终阻碍控制性能。为应对这些挑战,我们引入了HypEMBER,一种基于超网络和集成学习相结合的新型RL框架。在所提出的方法中,策略和价值函数都通过超网络来表示,这些超网络根据系统的物理参数生成基础模型的权重,从而实现跨不同动态范围的参数泛化。此外,使用策略和价值逼近器的集成来量化认知不确定性,从而在训练期间和之后改进探索策略并增强稳健性。在所提出框架的性能在两个代表性的参数化控制问题上进行了评估:(i)一维Kuramoto-Sivashinsky方程和(ii)二维时间相关涡旋流中的粒子导航任务,重点关注对测量噪声和参数错误指定的稳健性。数值结果表明,与现有RL方法相比,HypEMBER始终提高训练稳定性和样本效率,同时对影响系统动力学和可用观测的不确定性具有卓越的稳健性。
英文摘要
In this work we investigate reinforcement learning (RL) as a framework for the robust control of parametrized dynamical systems in presence of measurements and model uncertainties. High-dimensional state spaces, expensive numerical solvers, the partial knowledge of the governing equations, and the dependence on physical parameters that may be uncertain or difficult to estimate accurately, make the use of standard RL approaches computationally unfeasible. Indeed, lack of robustness and poor generalization across parameter variations are further amplified in presence of noisy or incomplete measurements, ultimately hampering control performance. To address these challenges, we introduce HypEMBER, a novel RL framework based on the combination of hypernetworks and ensemble learning. In the proposed approach, both the policy and value functions are represented through hypernetworks that generate the weights of the underlying models conditioned on the physical parameters of the system, thereby enabling parametric generalization across different dynamical regimes. In addition, an ensemble of policy and value approximators is employed to quantify epistemic uncertainty, leading to improved exploration strategies and enhanced robustness during and after training. The performance of the proposed framework is assessed on two representative parametrized control problems: (i) the one-dimensional Kuramoto-Sivashinsky equation and (ii) a particle-navigation task in a two-dimensional time-dependent gyre flow, focusing on robustness with respect to measurement noise and parameter misspecification. Numerical results demonstrate that HypEMBER consistently improves training stability and sample efficiency, while achieving superior robustness to uncertainties affecting both the system dynamics and the available observations, in comparison with state-of-the-art RL methods.