A Survey on Efficient Vision-Language-Action Models
高效视觉-语言-动作模型的综述
机构 * School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) ; School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机科学与人工智能学院) ; School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) ; Department of Information Engineering and Computer Science, University of Trento(特伦托大学信息工程与计算机科学系)
专题命中 VLA模型 :vision-language-action(title,abstract);action model(title,abstract);VLA(abstract);分类 cs.RO、cs.CV、cs.AI
AI总结 本文综述了高效视觉-语言-动作模型,系统分类了模型设计、训练和数据收集三个核心领域,总结了最新方法并提出了未来研究方向。
Comments 28 pages, 8 figures