CommentsAccepted by CVPR 2026 main conference. Compare to CVPR version, minor updates here are included (e.g., combine main text and appendix; clarify the timing scenario in appendix)
机构
*
School of Artificial Intelligence and Automation(人工智能与自动化学院)
;
Huazhong University of Science and Technology(华中科技大学)
;
School of Software(软件学院)
;
Henan University(河南大学)
;
Institute of Information Engineering(信息工程研究所)
;
Chinese Academy of Sciences(中国科学院)
;
School of Cyberspace Security(网络空间安全学院)
;
University of the Chinese Academy of Sciences(中国科学院大学)
;
Peking University(北京大学)
机构
*
Tencent Hunyuan(腾讯文元)
;
University of Maryland, College Park(马里兰大学 College Park 分校)
;
University of North Carolina, Chapel Hill(北卡罗来纳大学 Chapel Hill 分校)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
RubricEM: 通过规则引导的策略分解实现超越可验证奖励的元强化学习
Gaotang Li, Bhavana Dalvi Mishra, Zifeng Wang, Jun Yan, Yanfei Chen, Chun-Liang Li, Long T. Le, Rujun Han, George Lee, Hanghang Tong, Chen-Yu Lee, Tomas Pfister
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Google Cloud AI Research(谷歌云人工智能研究)
Policy Gradient Methods for Non-Markovian Reinforcement Learning
非马尔可夫强化学习中的策略梯度方法
Avik Kar, Siddharth Chandak, Rahul Singh, Soumitra Sinhahajari, Eric Moulines, Shalabh Bhatnagar, Nicholas Bambos
机构
*
Department of Computer Science and Automation, Indian Institute of Science(印度科学研究院计算机科学与自动化系)
;
Department of Electrical Engineering, Stanford University(斯坦福大学电气工程系)
;
Department of Electrical and Electronics Engineering, Nanyang Technological University(南洋理工大学电气与电子工程系)
;
CMAP, CNRS, École polytechnique, Institut Polytechnique d́e Paris(巴黎理工学院先进材料与工艺中心、国家科学研究中心、巴黎理工学院)
Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning
漂移是一种采样误差:面向长时间 horizon 的机器人规划的 SNR 意识功率分布
Kewei Chen, Yayu Long, Mingsheng Shang
机构
*
Chongqing Institute of Green and Intelligent Technology, Chinese Academy of Sciences(中国科学院重庆绿色智能技术研究院)
;
Chongqing School, University of Chinese Academy of Sciences(中国科学院大学重庆学院)
Safe and Real-Time Consistent Planning for Autonomous Vehicles in Partially Observed Environments via Parallel Consensus Optimization
通过并行一致性优化实现自动驾驶车辆在部分观测环境中的安全与实时一致性规划
Lei Zheng, Rui Yang, Minzhe Zheng, Michael Yu Wang, Jun Ma
机构
*
Robotics and Autonomous Systems Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)机器人与自主系统方向)
;
School of Engineering, Great Bay University(大湾大学工程学院)
;
Division of Emerging Interdisciplinary Areas, The Hong Kong University of Science and Technology(香港科学与技术大学新兴交叉领域学院)
ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning
ROS-LLM:一个用于具身AI的ROS框架,具有任务反馈和结构化推理
Christopher E. Mower, Yuhui Wan, Hongzhan Yu, Antoine Grosnit, Jonas Gonzalez-Billandon, Matthieu Zimmer, Jinlong Wang, Xinyu Zhang, Yao Zhao, Anbang Zhai, Puze Liu, Daniel Palenicek, Davide Tateo, Cesar Cadena, Marco Hutter, Jan Peters, Guangjian Tian, Yuzheng Zhuang, Kun Shao, Xingyue Quan, Jianye Hao, Jun Wang, Haitham Bou-Ammar
机构
*
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
University of Leeds(利兹大学)
;
Technical University of Darmstadt(达姆施塔特技术大学)
;
East China Normal University(华东师范大学)
;
Huawei Technologies(华为技术有限公司)
;
ETH Zurich(苏黎世联邦理工学院)
;
University College London(伦敦大学学院)