Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization
机构 * Xi’an Jiaotong University(西安交通大学) ; Westlake University(西湖大学) ; University of Texas at Austin(德克萨斯大学奥斯汀分校) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Georgia Institute of Technology(佐治亚理工学院) ; National University of Singapore(新加坡国立大学) ; Aalto University School of Electrical Engineering, Aalto University(阿alto大学电气工程学院,阿alto大学)