Habilis-$β$: A Fast-Motion and Long-Lasting On-Device Vision-Language-Action Model
Habilis-β:一种快速运动且持续运行的设备端视觉-语言-动作模型
Tommoro Robotics, :, Jesoon Kang, Taegeon Park, Jisu An, Soo Min Kimm, Jaejoon Kim, Jinu Pahk, Byungju Kim, Junseok Lee, Namheon Baek, Sungwan Ha, Hojun Baek, Eduardo Ayerve Cruz, Wontae Kim, Junghyeon Choi, Yousuk Lee, Joonmo Han, Sunghyun Cho, Sunghyun Kwon, Soyoung Lee, Jun Ki Lee, Seung-Joon Yi, Byoung-Tak Zhang, Theo Taeyeong Kim
Information-Theoretic Graph Fusion with Vision-Language-Action Model for Policy Reasoning and Dual Robotic Control
信息论图融合:基于视觉-语言-动作模型的政策推理与双臂机器人控制
Shunlei Li, Longsen Gao, Jin Wang, Chang Che, Xi Xiao, Jiuwen Cao, Yingbai Hu, Hamid Reza Karimi
机构
*
Electrical and Computer Engineering Department, University of New Mexico, Albuquerque, United States, 87106(电气与计算机工程系,新墨西哥大学,阿尔伯克基,美国,87106)
;
Dynamic Robot Systems Group, Oxford Robotics Institute, University of Oxford, United Kingdom, OX26NN(动态机器人系统组,牛津机器人研究所,牛津大学,英国,OX26NN)
;
Mechanical and Aerospace Engineering Department, The George Washington University, DC, United States, 22202(机械与航空航天工程系,乔治华盛顿大学,华盛顿特区,美国,22202)
;
Department of Computer Science, University of Alabama at Birmingham, Alabama, United States, 35294(计算机科学系,阿拉巴马大学伯明翰分校,阿拉巴马,美国,35294)
;
The School of Computation, Information and Technology, Technical University of Munich, Germany, 85748(计算、信息与技术学院,慕尼黑技术大学,德国,85748)
;
Department of Mechanical Engineering, Politecnico di Milano, Milan, Italy, 20156(机械工程系,米兰理工学院,米兰,意大利,20156)
机构
*
AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University(自动化实验室,人工智能学院,上海交通大学)
;
Anyverse Dynamics
;
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Terminal Technology Department, Alipay, Ant Group(终端技术部,蚂蚁集团)
Vision Language Action Models in Robotic Manipulation: A Systematic Review
视觉语言动作模型在机器人操作中的应用:系统综述
Muhayy Ud Din, Waseem Akram, Lyes Saad Saoud, Jan Rosell, Irfan Hussain
机构
*
Khalifa University Center for Autonomous Robotic Systems (KUCARS), Khalifa University, United Arab Emirates(卡利法大学自主机器人系统中心(KUCARS)、卡利法大学、阿拉伯联合酋长国)
;
Institute of Industrial and Control Engineering (IOC), Universitat Politecnica de Catalunya, Spain(工业与控制工程研究所(IOC)、巴塞罗那技术大学、西班牙)
专题命中
VLA模型
:vision language action(title,abstract);action model(title);VLA(abstract);分类 cs.RO、cs.CV
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
引导视觉-语言-动作模型作为反探索:一种测试时间缩放方法
Siyuan Yang, Yang Zhang, Haoran He, Ling Pan, Xiu Li, Chenjia Bai, Xuelong Li
机构
*
Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院)
;
University of Science and Technology of China(中国科学技术大学)
;
Tsinghua University(清华大学)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
基于注意力的视觉-语言-动作模型中逐块稀疏对抗攻击
Naifu Zhang, Wei Tao, Xi Xiao, Qianpu Sun, Yuxin Zheng, Wentao Mo, Peiqiang Wang, Nan Zhang
机构
*
Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(清华大学深圳国际研究生院)
;
Huazhong University of Science and Technology, Wuhan, China(华中科技大学)
;
Ping An Technology, Shenzhen, China(平安科技)
机构
*
Lanzhou University(兰州大学)
;
National University of Singapore(新加坡国立大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Tsinghua University(清华大学)
;
University of New South Wales(新南威尔士大学)
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
机构
*
Lanzhou University, China(兰州大学)
;
National University of Singapore, Singapore(新加坡国立大学)
;
Institute of Computing Technology, Chinese Academy of Sciences, China(中国科学院计算技术研究所)
;
NExT++ Research Centre, National University of Singapore, Singapore(新加坡国立大学NExT++研究中心)
专题命中
VLA模型
:vision language action(title,abstract);VLA(title,abstract);分类 cs.RO、cs.AI
Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
Haochuan Xu, Yun Sing Koh, Shuhuai Huang, Zirun Zhou, Di Wang, Jun Sakuma, Jingfeng Zhang
机构
*
The University of Auckland(奥克兰大学)
;
King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)
;
Tokyo University of Science(东京科学大学)
;
RIKEN Center for Advanced Intelligence Project(理化学研究所先进智能项目中心)