Internalizing World Models via Self-Play Finetuning for Agentic RL
Shiqi Chen, Tongyao Zhu, Zian Wang, Jinghan Zhang, Kangrui Wang, Siyang Gao, Teng Xiao, Yee Whye Teh, Junxian He, Manling Li
机构
*
City University of Hong Kong(香港城市大学)
;
Northwestern University(西北大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Oxford University(牛津大学)
;
Allen Institute for AI (AI2)(人工智能研究所)
;
University of Washington(华盛顿大学)
;
National University of Singapore(新加坡国立大学)
;
The Hong Kong Polytechnic University(香港理工大学)
Perfect Prediction or Plenty of Proposals? What Matters Most in Planning for Autonomous Driving
Aron Distelzweig, Faris Janjoš, Oliver Scheel, Sirish Reddy Varra, Raghu Rajan, Joschka Boedecker
机构
*
Department of Computer Science, University of Freiburg(弗赖堡大学计算机科学系)
;
Bosch Center for Artificial Intelligence(博世人工智能中心)
;
RWTH Aachen University(亚琛工业大学)
Pseudo-Kinematic Trajectory Control and Planning of Tracked Vehicles
Michele Focchi, Daniele Fontanelli, Davide Stocco, Riccardo Bussola, Luigi Palopoli
机构
*
Dipartimento di Ingegneria and Scienza dell'Informazione (DISI), University of Trento(信息工程系(DISI),特伦托大学)
;
Dipartimento di Ingegneria Industriale (DII), University of Trento(工业工程系(DII),特伦托大学)
Internet of Agents: Fundamentals, Applications, and Challenges
Yuntao Wang, Shaolong Guo, Yanghe Pan, Zhou Su, Fahao Chen, Tom H. Luan, Peng Li, Jiawen Kang, Dusit Niyato
机构
*
School of Cyber Science and Engineering, Xi'an Jiaotong University(网络安全科学与工程学院,西安交通大学)
;
School of Artificial Intelligence, Shandong University(人工智能学院,山东大学)
;
School of Automation, Guangdong University of Technology(自动化学院,广东技术大学)
;
College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)
Reinforcement Learning with Stochastic Reward Machines
Jan Corazza, Ivan Gavran, Daniel Neider
机构
*
University of Zagreb(Zagreb大学)
;
Max Planck Institute for Software Systems(软件系统马克斯·普朗克研究所)
专题命中
规划决策
:agent(abstract);分类 cs.AI、cs.LG
CommentsA shorter version of this paper appeared in the Proceedings of the Thirty-Sixth AAAI Conference on Artificial Intelligence (AAAI-22). Source code available at https://github.com/corazza/srm
Journal refCorazza, J., Gavran, I., & Neider, D. (2022). Reinforcement Learning with Stochastic Reward Machines. Proceedings of the AAAI Conference on Artificial Intelligence, 36(6), 6429-6436
Machine Learning-Based Ultrasonic Weld Characterization Using Hierarchical Wave Modeling and Diffusion-Driven Distribution Alignment
Joshua R. Tempelman, Adam J. Wachtor, Eric B. Flynn
机构
*
Data Science Group, Los Alamos National Laboratory, Los Alamos NM, USA(数据科学组,洛斯阿拉莫斯国家实验室)
;
Engineering Institute, Los Alamos National Laboratory, Los Alamos NM, USA(工程学院,洛斯阿拉莫斯国家实验室)
CLASP: General-Purpose Clothes Manipulation with Semantic Keypoints
Yuhong Deng, Chao Tang, Cunjun Yu, Linfeng Li, David Hsu
机构
*
School of Computing, Smart System Institute, National University of Singapore, Singapore(计算学院、智能系统研究所、新加坡国立大学,新加坡)
;
Department of Electronic and Electrical Engineering, Southern University of Science and Technology, China(电子与电气工程系、南方科技大学,中国)
Using matrices in post-processing phase of CFD simulations
Gianluca Argentini
专题命中
规划决策
:planning(abstract)
CommentsPaper based on presentation-talk at SCICOMP9, Bologna (Italy), March 23-26, 2004; workshop organized by IBM, CINECA (Italy) (dr. Sigismondo Boschi, dr. Giovanni Erbacci), NERSC-DOE (USA) (dr. David Skinner), web site: www.spscicomp.org ; main topics: Computational Fluid Dynamics
Journal refProgress in Industrial Mathematics at ECMI 2004 - Eindhoven (Netherlands), Springer, 2005