机构
*
Department of Electrical and Computer Engineering, University of California, Riverside, USA(电气与计算机工程系,加州大学河滨分校)
;
Thomas Lord Department of Computer Science, University of Southern California, USA(汤姆斯·劳德计算机科学系,南加州大学)
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
发现、学习与强化:通过多样化强化学习生成轨迹扩展视觉-语言-动作预训练
Rushuai Yang, Zhiyuan Feng, Tianxiang Zhang, Kaixin Wang, Chuheng Zhang, Li Zhao, Xiu Su, Yi Chen, Jiang Bian
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Tsinghua University(清华大学)
;
Wuhan University(武汉大学)
;
Central South University(中南大学)
;
Microsoft Research(微软研究院)
Path Planning through Multi-Agent Reinforcement Learning in Dynamic Environments
Jonas De Maeyer, Hossein Yarahmadi, Moharram Challenger
机构
*
Department of Computer Science University of Antwerp (UA)(安特卫普大学计算机科学系)
;
Department of Computer Engineering, Faculty of Engineering, Ayatollah Boroujerdi University(阿亚图拉·博鲁杰尔迪大学工程学院计算机工程系)
;
Department of Computer Science University of Antwerp (UA) and Flanders Make(安特卫普大学计算机科学系和弗拉芒制造)
DiAReL: Reinforcement Learning with Disturbance Awareness for Robust Sim2Real Policy Transfer in Robot Control
Mohammadhossein Malmir, Josip Josifovski, Noah Klarmann, Alois Knoll
机构
*
Department of Computer Engineering, School of Computation, Information and Technology, Technical University of Munich(计算机工程系,计算、信息与技术学院,慕尼黑技术大学)
;
Rosenheim University of Applied Sciences(罗森海姆应用技术大学)
专题命中
模仿学习与强化学习
:robotic(abstract);分类 cs.RO、cs.LG
CommentsAccepted for publication in IEEE Transactions on Control Systems Technology (TCST)
On the Convergence and Stability of Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers
Miroslav Štrupl, Oleg Szehr, Francesco Faccio, Dylan R. Ashley, Rupesh Kumar Srivastava, Jürgen Schmidhuber
机构
*
Dalle Molle Institute for Artificial Intelligence (IDSIA) - USI/SUPSI(达摩信息技术研究所(IDSIA)- USI/SUPSI)
;
Center of Excellence for Generative AI, King Abdullah University of Science and Technology(生成人工智能卓越中心,国王阿卜杜勒阿齐兹大学科学与技术学院)
;
NNAISENSE
机构
*
State Key Laboratory of Robotics and Intelligent Systems(机器人与智能系统国家重点实验室)
;
Shenyang Institute of Automation(沈阳自动化研究所)
;
Chinese Academy of Sciences(中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
Comments29 pages, 12 figures. Fazel Arasteh and Arian Haghparast contributed equally to this research. Submitted to ACM Transactions on Spatial Algorithms and Systems (TSAS). The code for this work is publicly available at https://github.com/Arianhgh/HHAN
机构
*
UC Berkeley(伯克利大学)
;
University of Oxford(牛津大学)
;
University of Washington(华盛顿大学)
;
UK AI Security Institute(英国人工智能安全研究所)
;
Google DeepMind(谷歌DeepMind)
Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm
Yang Chen, Menglin Zou, Jiaqi Zhang, Yitan Zhang, Junyi Yang, Gael Gendron, Libo Zhang, Jiamou Liu, Michael J. Witbrock
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
University of Auckland(奥克兰大学)
;
Chongqing University(重庆大学)
专题命中
模仿学习与强化学习
:robotics(abstract);分类 cs.AI、cs.LG
CommentsAccepted to NeurIPS 2025. Title used at submission and review: PIRO: Toward Stable Reward Learning for Inverse RL via Monotonic Policy Divergence Reduction
机构
*
Graduate School of Informatics, Nagoya University, Japan(名古屋大学信息学研究科)
;
RIKEN Center for Advanced Intelligence Project, Japan(RIKEN高级智能项目研究中心)
;
Graduate School of Arts and Sciences, The University of Tokyo, Japan(东京大学文学系研究科)
;
Project team for SIP, Japan(SIP项目团队)
;
Japan Agency for Marine-Earth Science and Technology, Japan(日本海洋地球科学技术机构)
;
Faculty of Education, Shitennoji University, Japan(世田谷大学教育学部)
;
Graduate School of Engineering, The University of Tokyo, Japan(东京大学工学研究科)
;
Graduate School of Science and Technology, Niigata University, Japan(新潟大学科学技术研究科)
;
Graduate School of Science, Nagoya University, Japan(名古屋大学理学研究科)
;
Principles of Informatics Research Division, National Institute of Informatics, Japan(信息学原理研究部门,日本信息处理技术研究所)
;
Graduate School of Information Science, The University of Osaka, Japan(大阪大学信息科学研究科)