arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24603cs.RO

感知夹具的视觉语言动作模型

Gripper-aware Vision Language Action Models

  • University of Liverpool(利物浦大学)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • Indian Institute of Science(印度科学学院)
  • The University of Tokyo(东京大学)
  • Zürcher Hochschule für Angewandte Wissenschaften(苏黎世应用科技大学)
  • University of Arkansas(阿肯色大学)
  • Physical Intelligence(物理智能公司)
  • Huazhong University of Science and Technology(华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Hanyi Zhang, Zihong Luo, Tianyu Li, Khang Nguyen, Basu Hela, Shreyas Kumar, Ngoc Duy Tran, Feng Dai, Charith Munasinghe, Jorge Peña Queralta, Giovanni Toffetti,… 展开作者

Hanyi Zhang, Zihong Luo, Tianyu Li, Khang Nguyen, Basu Hela, Shreyas Kumar, Ngoc Duy Tran, Feng Dai, Charith Munasinghe, Jorge Peña Queralta, Giovanni Toffetti, Khoa Vo, Ngan Le, Ravi Prakash, Quan Vuong, Tung D. Ta, Long Hu, Anh Nguyen, Baoru Huang

AI总结:

针对现有视觉语言动作模型(VLA)忽略夹具差异的问题,提出多夹具感知数据集MiGA与结合多夹具分词器及适配器策略路由的GVLA,实验显示其性能优于基线且泛化与适应能力更强。

AI中文摘要:

视觉语言动作模型(VLAs)通过让机器人解读视觉观测结果与自然语言指令来生成可执行的动作序列,已在通用机器人抓取与操作任务中取得进展。但现有VLAs常默认夹具不变,而抓取策略本质上依赖于机器人本体,不同夹具类型(如平行夹爪与吸盘)为达成相同抓取目标通常需要不同的交互策略。此外,当前VLAs数据集大多依赖平行夹爪,限制了感知夹具的学习。为解决该缺口,我们引入MiGA,这是一个涵盖五种不同夹具类型、跨多个机器人、包含103000次演示的多夹具感知数据集,明确捕捉相同任务目标下的策略差异。我们还提出GVLA,它结合了新型多夹具分词器与基于适配器的策略路由。我们的新型夹具编码能生成结构化嵌入信息,平衡参数共享与策略区分,而分层探测证实了VLAs中具有意义的夹具条件表征。在仿真与真实机器人上的大量实验显示,我们的GVLA在所有评估设置下均优于当前基线方法,还提升了对新物体或未见过任务的零样本泛化或少样本适应能力,并实现了更高效的夹具适应。

英文摘要:

Vision language action models (VLAs) have advanced general purpose robotic grasping and manipulation by enabling robots to interpret visual observations and natural language instructions to generate executable action sequences. However, existing VLAs often implicitly assume gripper invariance, despite grasping strategies being inherently embodiment-dependent. Different gripper types, such as parallel-jaw and suction, usually require distinct interaction strategies to achieve the same grasping objective. Moreover, current datasets for VLAs predominantly rely on parallel-jaw grippers, limiting gripper-aware learning. To address this gap, we introduce MiGA, a multi-gripper-aware dataset spanning five distinct gripper types across multiple robots with 103,000 demonstrations, explicitly capturing strategy divergence under shared task objectives. We further propose GVLA, which combines a new multi-gripper tokenizer with adapter-based policy routing. Our new gripper encoding induces structured embedding information that balances parameter sharing and strategy differentiation, while layer-wise probing confirms meaningful gripper-conditioned representations for VLAs. Intensive experiments in both simulation and real-world robots show that our GVLA outperforms the current baselines across evaluated settings. Our method also improves zero-shot generalization or few-shot adaptation to new objects or unseen tasks, and enable more efficient gripper adaptation.

↑