CommentsThe results presented in this paper are preliminary. Please note that the experiments are currently ongoing, and the final data is subject to change upon the completion of the study. All ideas, results, methods, and any content herein are the sole property of the authors
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
See-Control: 一种多模态代理框架用于智能手机与机械臂的交互
Haoyu Zhao, Weizhong Ding, Yuhao Yang, Zheng Tian, Linyi Yang, Kun Shao, Jun Wang
机构
*
University College London(伦敦大学学院)
;
Imperial College London(伦敦帝国学院)
;
Huawei Noah’s Ark Lab(华为诺亚实验室)
;
ShanghaiTech University(上海科技大学)
;
Southern University of Science(南方科技大学)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
A Survey on Vision-Language-Action Models for Autonomous Driving
Sicong Jiang, Zilin Huang, Kangan Qian, Ziang Luo, Tianze Zhu, Yang Zhong, Yihong Tang, Menglin Kong, Yunlong Wang, Siwen Jiao, Hao Ye, Zihao Sheng, Xin Zhao, Tuopu Wen, Zheng Fu, Sikai Chen, Kun Jiang, Diange Yang, Seongjin Choi, Lijun Sun
机构
*
McGill University(麦吉尔大学)
;
Tsinghua University(清华大学)
;
Xiaomi Corporation(小米公司)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
University of Minnesota–Twin Cities(明尼苏达大学双城分校)
;
State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University(智能绿色车辆与移动国家重点实验室,清华大学)
专题命中
GUI与屏幕智能体
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI