arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25864cs.RO

MA-VLA:用于协作与组合泛化的多臂视觉-语言-动作模型

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

Zaibin Zhang, Junlan Xiao, Zhongbo Zhang, Yifan Wang, Li Kang, Yiran Qin, Changxing Xia, Heng Zhou, Talas Fu, Enshen Zhou, Ruimao Zhang, Zhenfei Yin, Huchuan Lu, Lijun Wang

首次发表
浏览论文内容

中文总结 AI 辅助

MA-VLA通过将协作行为分解为原子提示并引入Arm Shuffle训练方法,解决了多臂VLA模型的协作泛化问题,在未见过的协作模式下表现优于现有模型。

中文摘要 AI 辅助

多臂协作正成为具身操控的核心能力。近期的视觉-语言-动作(VLA)模型整合了感知、语言与控制,但多数模型将语言表示为单一全局指令,未提供明确的机制来分配和组合各臂的特定行为,这种设计限制了其向训练期间未观察到的协作模式的迁移。我们提出MA-VLA,这是一种通过原子动作分配实现多臂协作的统一框架。MA-VLA将协作行为分解为中级原子提示,并将其分配给各臂,从而实现明确的子目标指定和跨任务的组合复用。为减少对固定执行角色的依赖,我们引入了Arm Shuffle,这是一种训练时对各臂的观测、状态及分配的原子提示进行置换的方法,该置换强化了角色无关的指令遵循,并支持重组为未见过的协调模式,我们将此称为多臂组合泛化。我们还构建了一个基准,其中测试时的协作模式在训练集中不存在。在仿真和真实世界评估中,现有最先进的VLA模型在这些未见过的协作场景中大多失败,而MA-VLA始终能成功。这些结果表明,结构化的每臂原子动作分配为多臂具身系统的可扩展泛化提供了可行途径。代码、模型和数据可在https URL获取。

英文摘要

Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) models integrate perception, language, and control, but most represent language as a single global instruction and do not provide an explicit mechanism for assigning and composing arm-specific behaviors. This design limits transfer to collaboration patterns that differ from those observed during training. We present MA-VLA, a unified framework for multi-arm collaboration via atomic action assignment. MA-VLA decomposes cooperative behavior into mid-level atomic prompts and allocates them to individual arms, enabling explicit subgoal specification and compositional reuse across tasks. To reduce reliance on fixed execution roles, we introduce Arm Shuffle, a training-time permutation of the observation, state, and assigned atomic prompts for each arm. This permutation enforces role-agnostic instruction following and supports recomposition into unseen coordination patterns, which we term multi-arm compositional generalization. We also construct a benchmark in which test-time collaboration patterns are absent in training set. Across simulation and real-world evaluations, prior state-of-the-art VLAs largely fail under these unseen collaborations, while MA-VLA consistently succeeds. These results indicate that structured, per-arm atomic action assignment offers a practical route to scalable generalization in multi-arm embodied systems. Code, models, and data are available at https://github.com/zhangzaibin/future-robots

发表机构

  • Dalian University of Technology(大连理工大学)
  • University of Oxford(牛津大学)
  • Sun Yat-sen University(中山大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • University of Science and Technology of China(中国科学技术大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑