Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning
机构 * Department of Mathematical Sciences Tsinghua University(清华大学数学科学系) ; School of Vehicle and Mobility Tsinghua University(清华大学车辆与移动系统学院) ; Department of Mathematical Sciences, Tsinghua University(清华大学数学科学系) ; School of Vehicle and Mobility & College of AI Tsinghua University(清华大学车辆与移动系统学院与人工智能学院)
Comments Submitted to the Journal of Machine Learning Research (JMLR), under review