Comments10 main content pages, 4 main content figures, 11 appendix pages, 5 appendix figures, camera ready version submitted to 39th Conference on Neural Information Processing Systems (NeurIPS 2025)
Sample-Efficient Expert Query Control in Active Imitation Learning via Conformal Prediction
通过置信预测实现高效的专家查询控制在主动模仿学习中
Arad Firouzkouhi, Omid Mirzaeedodangeh, Lars Lindemann
机构
*
Department of Computer Science, University of Southern California(计算机科学系,南加州大学)
;
Department of Information Technology and Electrical Engineering, ETH Zurich(信息科技与电气工程系,苏黎世联邦理工学院)
Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended Environments
深度强化学习需要深度行为分析:通过无模型智能体在开放性环境中探索隐式规划
Riley Simmons-Edler, Ryan P. Badman, Felix Baastad Berg, Raymond Chua, John J. Vastola, Joshua Lunger, William Qian, Kanaka Rajan
机构
*
Department of Neurobiology, Harvard Medical School(哈佛医学院神经生物学系)
;
Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究学院)
;
Department of Mathematics, NTNU(NTNU数学系)
;
School of Computer Science, McGill University & Mila(麦吉尔大学计算机科学学院及Mila)
;
Department of Computer Science, University of Toronto(多伦多大学计算机科学系)
;
Biophysics Graduate Program, Harvard University(哈佛大学生物物理学研究生项目)
机构
*
Purdue University(普渡大学)
;
Beijing University of Chemical Technology(北京化工大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Sungkyunkwan University(成均馆大学)
;
Indiana University Bloomington(印第安纳大学布卢明顿分校)
On the Tension Between Optimality and Adversarial Robustness in Policy Optimization
在策略优化中最优性与对抗鲁棒性之间的张力
Haoran Li, Jiayu Lv, Congying Han, Zicheng Zhang, Anqi Li, Yan Liu, Tiande Guo, Nan Jiang
机构
*
School of Mathematical Sciences University of Chinese Academy of Sciences(中国科学院大学数学科学学院)
;
School of Mathematical Sciences Nankai University(南开大学数学科学学院)
;
School of Statistics and Data Science Nankai University(南开大学统计与数据科学学院)
;
Siebel School of Computing and Data Science University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校Siebel计算与数据科学学院)
专题命中
模仿学习与强化学习
:navigation(abstract);分类 cs.LG
AI总结
本文提出 BARPO 框架,通过调节对手强度统一 SPO 和 ARPO,解决策略优化中最优性与对抗鲁棒性之间的张力问题。
机构
*
College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院)
;
Zhejiang University-University of Illinois Urbana-Champaign (ZJU-UIUC) Institute, Zhejiang University(浙江大学-伊利诺伊大学厄巴纳-香槟分校联合研究所)
;
Department of Engineering, King’s College London(伦敦国王学院工程系)