arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

双臂视觉-语言-动作模型中的臂级组合泛化

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models

Zaibin Zhang, Binghao Ran, Yuhan Wu, Zhongbo Zhang, Yifan Wang, Junwei Jiang, Junlan Xiao, Wangcheng Shi, Li Kang, Yiran Qin, Zhenfei Yin, Lijun Wang, Huchuan Lu

arXiv 2610.06184首次发表:更新:

AI 中文总结

本文提出ACG-Bench基准,研究双臂策略的臂级组合泛化,并构建AE-VLA模型,在模拟和物理机器人上显著提升泛化成功率。

AI 中文摘要

多臂协作中的泛化可以被视为以新的方式跨臂组合熟悉的原子技能。然而,现有的评估对于在不同协调要求下哪些训练和架构选择支持这种能力提供的见解有限。我们引入了ACG-Bench,一个用于臂级组合泛化的基准,为研究双臂策略中的技能重组提供了一个通用测试平台。它包含8个任务族中的23个任务-条件对,其中6个为域内条件,17个为未见过的组合,涵盖重排序、同步、它们的组合以及跨任务组合。所有方法都接收相同的每臂原子提示,成功要求是在满足物理里程碑和指定的顺序或时间约束的同时实现任务目标。使用π_{0.5}作为共同的视觉-语言-动作骨干,我们比较了具有共享源数据和共同评估协议的代表性数据增强和架构策略。我们的架构研究考察了臂标记分组、技能特定的LoRA适配器(SkillLoRA)和臂级注意力(AWA),强调了技能条件参数和注意力结构的互补性。将这些选择组合起来得到AE-VLA,其在模拟中实现了21.53%的泛化成功率,而单个π_{0.5}为2.94%,MA-VLA为3.06%,两个独立控制的π_{0.5}策略为5.53%。在物理SO101机器人上,AE-VLA在五个未见条件下达到39.00%的平均成功率,而最强基线为10.00%。这些发现为设计超越固定训练例程的双臂策略提供了经验指导。

英文摘要

Generalization in multi-arm collaboration can be studied as composing familiar atomic skills in new ways across arms. However, existing evaluations offer limited insight into which training and architectural choices support this ability under different coordination requirements. We introduce \textbf{ACG-Bench}, a benchmark for \emph{Arm-wise Compositional Generalization} that provides a common testbed for studying skill recomposition in dual-arm policies. It contains 23 task--condition pairs across 8 task families, with 6 in-domain conditions and 17 unseen compositions covering reordering, synchronization, their combination, and cross-task composition. All methods receive the same per-arm atomic prompts, and success requires achieving the task goal while satisfying physical milestones and specified order or timing constraints. Using $π_{0.5}$ as a common vision-language-action backbone, we compare representative data-augmentation and architectural strategies with shared source data and a common evaluation protocol. Our architectural study examines arm-token grouping, skill-specific LoRA adapters (SkillLoRA), and arm-wise attention (AWA), highlighting the complementarity of skill-conditioned parameters and attention structure. Combining these choices yields \textbf{AE-VLA}, which achieves 21.53\% generalization success in simulation, compared with 2.94\% for Single $π_{0.5}$, 3.06\% for MA-VLA, and 5.53\% for two independently controlled $π_{0.5}$ policies. On physical SO101 robots, AE-VLA reaches 39.00\% mean success across five unseen conditions, compared with 10.00\% for the strongest baseline. These findings provide empirical guidance for designing dual-arm policies that generalize beyond fixed training routines.

CommentsTechnical Report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑