arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16300eess.SYcs.SY

基于历史依赖策略类别的策略梯度方法用于域随机化线性二次调节器

Policy Gradient over History-Dependent Policy Classes for LQR with Domain Randomization

  • University of Pennsylvania(宾夕法尼亚大学)
  • ETH AI Center(苏黎世联邦理工学院人工智能中心)
  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Tesshu Fujinami, Bruce D. Lee, Anastasios Tsiamis, Nikolai Matni, George J. Pappas

AI总结:

本研究针对域随机化LQR问题,提出基于历史依赖策略类别(如有限脉冲响应控制器)的策略梯度方法,并设计课程学习算法逐步扩展控制器记忆,证明在环境异质性有界时全局收敛,实验验证了方法的有效性。

AI中文摘要:

域随机化(DR)已被广泛用于通过强化学习在模拟环境分布上训练控制器,以克服模拟到现实的差距。虽然DR仅使用通过策略梯度(PG)方法合成的控制器即可实现鲁棒性能,但其优化景观尚未被充分理解,即使在线性二次调节器(LQR)目标的情况下也是如此。为此,我们首先研究域随机化LQR在历史依赖策略类别(如有限脉冲响应控制器)上的策略梯度,因为这些控制器可以扩展同时稳定的可能性。其次,为了找到这样的稳定控制器,我们提出了一种基于课程学习的算法,该算法逐步扩展控制器的记忆。最后,我们证明在环境异质性适当有界的条件下,所提出的算法结合策略梯度能够全局收敛到DR目标样本平均近似的最小化器。实验结果支持我们的发现,并突出了未来工作的有前景方向,包括非线性域随机化控制。

英文摘要:

Domain Randomization (DR) has been widely used to overcome the sim-to-real gap by training a controller on a distribution of simulated environments via reinforcement learning. While DR can achieve robust performance simply using controllers synthesized via policy gradient (PG) methods, the optimization landscape is not well understood, even in the case of linear quadratic regulator (LQR) objectives. To this end, we first study PG of domain randomized LQR over history-dependent policy classes, such as finite impulse response controllers, as they can extend the possibilities of simultaneous stabilization. Second, to find such a stabilizing controller, we propose a curriculum learning based algorithm which gradually expands the memory of the controller. Finally, we show that PG with the proposed algorithm converges globally to the minimizer of a sample average approximation of the DR objective under suitable bounds on the heterogeneity of environments. Empirical results support our findings and highlight promising directions for future work, including nonlinear domain-randomized control.

↑