Gripper-aware Vision Language Action Models
感知夹具的视觉语言动作模型
机构 * University of Liverpool(利物浦大学) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) ; Indian Institute of Science(印度科学学院) ; The University of Tokyo(东京大学) ; Zürcher Hochschule für Angewandte Wissenschaften(苏黎世应用科技大学) ; University of Arkansas(阿肯色大学) ; Physical Intelligence(物理智能公司) ; Huazhong University of Science and Technology(华中科技大学)
AI总结 针对现有视觉语言动作模型(VLA)忽略夹具差异的问题,提出多夹具感知数据集MiGA与结合多夹具分词器及适配器策略路由的GVLA,实验显示其性能优于基线且泛化与适应能力更强。