机构
*
The Chinese University of Hong Kong(香港中文大学)
;
The Third Affiliated Hospital of Sun Yat-sen University(中山大学附属第三医院)
;
University of Cambridge(剑桥大学)
;
University of California, Davis(加州大学戴维斯分校)
Ye Yuan, Kehan Chen, Xinqiang Yu, Wentao Xu, Heng Wang, Libo Huang, Chuanguang Yang, Yan Huang, Jiawei He, Zhulin An
机构
*
School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所人工智能安全国家重点实验室)
;
National Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
XYZ Embodied AI(XYZ具身人工智能)
机构
*
College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)
;
Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院智能信息处理上海市重点实验室)
;
TeleAI, China Telecom(中国电信天翼人工智能公司)
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
XR-1:通过学习统一的视觉-运动表示实现多功能的视觉-语言-动作模型
Shichao Fan, Kun Wu, Zhengping Che, Xinhua Wang, Di Wu, Fei Liao, Ning Liu, Yixue Zhang, Zhen Zhao, Zhiyuan Xu, Meng Li, Qingjie Liu, Shanghang Zhang, Min Wan, Jian Tang
机构
*
Beijing Innovation Center of Humanoid Robotics, Beijing, China(北京人形机器人创新中心,北京,中国)
;
School of Mechanical Engineering and Automation, Beihang University, Beijing, China(北京航空航天大学机械工程及自动化学院,北京,中国)
;
State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,SCSE,北京航空航天大学,北京,中国)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)
机构
*
Department of Artificial Intelligence, School of Informatics, Xiamen University(厦门大学信息学院人工智能系)
;
Department of Computer Science, Aberystwyth University(阿伯里斯特威斯大学计算机科学系)
CommentsWe have further refined the benchmark construction and experimental presentation to improve clarity and consistency. The revised version includes updated task design, food-resource data, and evaluation details to better align the benchmark with the intended food resource referral setting. These changes provide a more precise presentation of the experimental findings
Bridging the Agent-World Gap: Text World Models for LLM-based Agents
弥合智能体-世界鸿沟:面向基于LLM的智能体的文本世界模型
Yixia Li, Hongru Wang, Peng Lai, Zhiwen Ruan, He Zhu, Youxin Zhu, Ganlong Zhao, Minda Hu, Yun Chen, Sibei Yang, Peng Li, Jeff Z. Pan, Jia Pan, Guanhua Chen, Yang Liu, Guanbin Li
机构
*
Southern University of Science and Technology(南方科技大学)
;
University of Edinburgh(爱丁堡大学)
;
Peking University(北京大学)
;
Sun Yat-sen University(中山大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai University of Finance and Economics(上海财经大学)
;
Tsinghua University(清华大学)
;
The University of Hong Kong(香港大学)
机构
*
Pengcheng Laboratory(鹏城实验室)
;
School of Computer Science and Cyber Engineering(计算机科学与网络工程学院)
;
Guangzhou University(广州大学)
;
Southern University of Science and Technology(南方科技大学)
机构
*
Department of Computer Science and Engineering, Indian Institute of Technology Patna, India(印度理工学院帕纳瓦分校计算机科学与工程系)
;
School of Information Technology, King Mongkut’s Institute of Technology Ladkrabang, Thailand(泰国拉差班国王理工大学信息科技学院)
机构
*
University of Chinese Academy of Sciences(中国科学院大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Tsinghua University(清华大学)
;
Zhejiang University(浙江大学)
;
Beihang University(北京航空航天大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
The University of Hong Kong(香港大学)
机构
*
School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
;
Meituan(美团)
;
Zhejiang University(浙江大学)
;
The Chinese University of Hong Kong(香港中文大学)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract)
RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models
RoVLA: 多一致性约束用于鲁棒的视觉-语言-动作模型
Jingzhou Luo, Yifan Wen, Yongjie Bai, Xinshuai Song, Yang Liu, Liang Lin
机构
*
Sun Yat-sen University(中山大学)
;
Peng Cheng Laboratory(鹏城实验室)
;
Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)
;
X-Era AI Lab(X-Era AI实验室)
机构
*
College of Information Science and Technology, Beijing University of Chemical Technology(北京化工大学信息科学与技术学院)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
Institute of Systems Engineering and Collaborative Laboratory for Intelligent Science and Systems, Macau University of Science and Technology(澳门大学系统工程与智能科学与系统联合实验室)
;
School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)