Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning
半无限安全强化学习的交换策略优化算法
机构 * Department of Mathematical Sciences Tsinghua University(清华大学数学科学系) ; School of Vehicle and Mobility Tsinghua University(清华大学车辆与移动系统学院) ; Department of Mathematical Sciences, Tsinghua University(清华大学数学科学系) ; School of Vehicle and Mobility & College of AI Tsinghua University(清华大学车辆与移动系统学院与人工智能学院)
AI总结 针对半无限安全强化学习的无限约束难题,提出交换策略优化算法,可提供有界安全保证并收敛至最优策略,推导了迭代次数上界与最优值差距。
Comments Submitted to the Journal of Machine Learning Research (JMLR), added new experiments and expanded analysis