PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
面向通用视觉-语言-动作策略的通用姿态预训练
Haitao Lin, Hanyang Yu, Jingshun Huang, He Zhang, Yonggen Ling, Ping Tan, Xiangyang Xue, Yanwei Fu
机构
*
Tencent Robotics X(腾讯机器人X)
;
Futian Laboratory(福田实验室)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Fudan University(复旦大学)
;
Shanghai Innovation Institute(上海创新研究院)
机构
*
School of Mathematics and Statistics and Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(数学与统计学学院和教育部智能网络与网络安全重点实验室,西安交通大学)
;
Tencent AI Lab(腾讯AI实验室)
;
ByteDance(字节跳动)
;
SINGPATH AI Lab, SingPath Medical Technology Pte. Ltd.(SingPath AI实验室,SingPath医疗科技私人有限公司)
;
School of Mathematics and Statistics, Xidian University(数学与统计学学院,西安电子科技大学)
;
College of Biomedical Engineering, Sichuan University(生物医学工程学院,四川大学)
;
School of Mathematics and Statistics, Ministry of Education Key Lab of Intelligent Networks and Network Security, Xi’an Jiaotong University(数学与统计学学院和教育部智能网络与网络安全重点实验室,西安交通大学)
;
Macao Institute of Systems Engineering, Macau University of Science and Technology(澳门系统工程研究院,澳门科技大学)