VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction
VANE:通过未来视觉表示预测实现视觉-语言-动作模型的可靠测试时训练
Hongjin Ji, Guoyang Xia, Luoyang Sun, Fangxiang Feng, Lei Ren
机构
*
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Li Auto Inc.(理想汽车)
机构
*
Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)
;
Department of Engineering Science, University of South Florida(工程科学系,佛罗里达州立大学)
;
Department of Computer Science, Rutgers University(计算机科学系,罗格斯大学)
RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models
RLRC:基于强化学习的压缩视觉-语言-动作模型恢复
Yuxuan Chen, Yixin Han, Yize Huang, Xiao Li
机构
*
State Key Laboratory of Mechanical System and Vibration(机械系统与振动国家重点实验室)
;
Shanghai Key Laboratory of Intelligent Robotics(上海智能机器人重点实验室)
;
School of Mechanical Engineering, Shanghai Jiao Tong University(上海交通大学机械工程学院)
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
Evo-Depth:一种轻量化的深度增强视觉-语言-动作模型
Tao Lin, Yuxin Du, Jiting Liu, Nuobei Zhu, Yunhe Li, Yuqian Fu, Yinxinyu Chen, Hongyi Cai, Zewei Ye, Bing Cheng, Kai Ye, Yiran Mao, Yilei Zhong, MingKang Dong, Junchi Yan, Gen Li, Bo Zhao
机构
*
School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院)
;
King Abdullah University of Science and Technology(卡塔尔国王 Abdullah 大学科学与技术大学)
;
Nanyang Technological University(南洋理工大学)
;
SJTU-Quic Robot Joint Lab(上海交通大学-Quick 机器人联合实验室)
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
SafeVLA: 通过约束学习实现视觉-语言-动作模型的安全对齐
Borong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei, Yishuai Cai, Josef Dai, Yuanpei Chen, Yaodong Yang
机构
*
Institute for Artificial Intelligence, Peking University(人工智能研究院,北京大学)
;
PKU-PsiBot Joint Lab(北京大学PsiBot联合实验室)
;
State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学)
;
Zhongguancun Academy(中关村学院)
Yitao Jiang, Roy Xing, Luyang Zhao, Brian Plancher, Muhao Chen, Devin Balkcom
机构
*
Dept. of Computer Science, Dartmouth College(达特茅斯学院计算机科学系)
;
Dept. of Electrical and Computer Engineering, Clemson University(克莱姆森大学电气与计算机工程系)
;
Dept. of Mechanical and Aerospace Engineering, University of Houston(休斯顿大学机械与航空航天工程系)
CommentsThe results presented in this paper are preliminary. Please note that the experiments are currently ongoing, and the final data is subject to change upon the completion of the study. All ideas, results, methods, and any content herein are the sole property of the authors
机构
*
MoE Key Lab of Artificial Intelligence, Institute of AI, Shanghai Jiao Tong University(教育部人工智能重点实验室,上海交通大学人工智能研究院)
;
Central Research Institute, Huawei(华为中央研究院)
机构
*
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
Beijing Academy of Artificial Intelligence (BAAI)(北京智源人工智能研究院)
;
Beihang University(北京航空航天大学)
;
Eastern Institute of Technology, Ningbo(宁波东方理工大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Microsoft Research Asia (MSRA)(微软亚洲研究院)
专题命中
VLA模型
:vision-language-action(title,abstract);action model(title);VLA(abstract,abstract_cn);embodied foundation model(abstract)