RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models
RewardFlow: 面向大语言模型智能体强化学习的拓扑感知状态图奖励传播
Xiao Feng, Bo Han, Zhanke Zhou, Jiaqi Fan, Jiangchao Yao, Ka Ho Li, Dahai Yu, Michael Kwok-Po Ng
机构
*
TMLR Group(TMLR小组)
;
Hong Kong Baptist University(香港 Baptist大学)
;
TCL Corporate Research (HK) Co Ltd(TCL企业研究(香港)有限公司)
;
Cooperative Medianet Innovation Center Shanghai Jiao Tong University(合作中位网创新中心上海交通大学)
;
Department of Mathematics Hong Kong Baptist University(香港 Baptist大学数学系)
机构
*
Center for Advanced Robotics Technology Innovation (CARTIN), School of Electrical and Electronic Engineering, Nanyang Technological University(先进机器人技术与创新中心(CARTIN)、电子工程学院、南洋理工大学)
;
Spatial AI & Robotics Lab, Department of Computer Science and Engineering(空间人工智能与机器人实验室、计算机科学与工程系)
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment
通过几何奖励信用分配强化点-视觉-语言模型的3D理解
Jingkun Chen, Ruoshi Xu, Mingqi Gao, Shengda Luo, Jungong Han
机构
*
Northwestern Polytechnical University(西北工业大学)
;
Southern University of Science and Technology(南方科技大学)
;
The University of Sheffield(谢菲尔德大学)
;
Hengqin Laboratory(横琴实验室)
;
Tsinghua University(清华大学)
Yuwei Ning, Ganlong Zhao, Yipeng Qin, Si Liu, Yang Liu, Liang Lin, Guanbin Li
机构
*
Sun Yat-sen University(中山大学)
;
Peng Cheng Laboratory(鹏城实验室)
;
The Chinese University of Hong Kong(香港中文大学)
;
Centre for Perceptual and Interactive Intelligence(感知与交互智能中心)
;
Cardiff University(卡迪夫大学)
;
Beihang University(北航)
;
Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)
Hang Ye, Xiaoxuan Ma, Fan Lu, Wayne Wu, Kwan-Yee Lin, Yizhou Wang
机构
*
Peking University(北京大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Tongji University(同济大学)
;
University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
;
University of Michigan(密歇根大学)
机构
*
National Yang Ming Chiao Tung University(国立阳明交通大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
AI Research Center, Hon Hai Research Institute(鸿海研究院人工智能研究中心)
机构
*
Bioengineering Department, Imperial College London, London, UK School of Biomedical Engineering \& lmaging Sciences, King's College London, London,UK
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
SAMoE-VLA:一种面向自动驾驶的场景自适应混合专家视觉-语言-动作模型
Zihan You, Hongwei Liu, Chenxu Dang, Zhe Wang, Sining Ang, Aoqi Wang, Yan Wang
机构
*
Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学)
;
School of Instrument Science and Engineering, Southeast University(仪器科学与工程学院,东南大学)
;
Zhili College, Tsinghua University(紫荆学院,清华大学)
;
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学)
;
Department of Automation, University of Science and Technology of China(自动化学院,中国科学技术大学)
;
Department of Automation, University of Science and Technology Beijing(自动化学院,北京科技大学)
T2Nav Algebraic Topology Aware Temporal Graph Memory and Loop Detection for ZeroShot Visual Navigation
T2Nav 基于代数拓扑的时序图记忆与循环检测用于零样本视觉导航
Quang-Anh N. D., Duc Pham, Minh-Anh Nguyen, Tung Doan, Tuan Dang
机构
*
International School, Vietnam National University, Hanoi, Vietnam(越南国家大学河内国际学校)
;
Hanoi University of Science, Vietnam National University, Hanoi, Vietnam(越南国家大学河内科学大学)
;
Hanoi University of Science and Technology, Hanoi, Vietnam(河内科学技术大学)
;
Cognitive Robotics Lab, Department of Electrical Engineering and Computer Science, University of Arkansas, Fayetteville, AR, USA(美国阿肯色大学富尔顿分校电气工程与计算机科学系认知机器人实验室)
;
University of Arkansas, Fayetteville, AR, USA(美国阿肯色大学富尔顿分校)