ZETA:面向桌面操作的零样本跨 embodiment VLA 迁移的受控研究
ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation
浏览论文内容
中文总结 AI 辅助
ZETA研究区分两类零样本跨embodiment VLA迁移,引入含14个目标embodiment的受控基准,发现状态-动作表示等因素可提升迁移性能,为桌面操作场景提供实用指导。
中文摘要 AI 辅助
对未见 embodiment 的零样本泛化对于可泛化的视觉-语言-动作(VLA)模型至关重要,因为机器人硬件不断发展,且特定任务的数据采集成本高昂。然而,对该问题的系统理解仍然有限,部分原因是文献缺乏统一的零样本迁移定义,以及能将 embodiment 变化与任务、场景或协议差异隔离开的受控评估设置。为解决这一差距,我们首先区分严格零样本迁移(目标 embodiment 未出现在所有训练数据中)与预训练暴露型零样本迁移(目标 embodiment 仅在预训练期间出现)。随后,我们引入了一个受控基准,涵盖模拟和真实世界验证中的14个保留目标 embodiment。在该框架内,我们对四个因素进行了受控分析:状态-动作表示、预训练 embodiment 多样性、辅助联合训练目标以及目标 embodiment 暴露。实验结果表明,局部末端执行器(EEF)状态-动作表示、源 embodiment 多样性以及辅助联合训练分别将跨 embodiment 迁移性能提升了约15、18和7个百分点。我们进一步发现,在预训练期间仅添加5%的目标 embodiment 数据,可使平均目标 embodiment 进展提升13.4个百分点,这表明严格零样本迁移与预训练暴露型零样本迁移是不同的,应分别报告。这些发现共同为评估和改进两指夹爪桌面操作场景中的跨 embodiment VLA 迁移提供了实用指导,同时推动未来对更广泛设置的研究,包括移动基座控制、灵巧手和长 horizon 任务。
英文摘要
Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) models as robot hardware evolves and task-specific data collection remains costly. However, a systematic understanding of this problem remains limited, in part because the literature lacks a unified zero-shot transfer definition and controlled evaluation settings that isolate embodiment changes from differences in tasks, scenes, or protocols. To address this gap, we first distinguish strict zero-shot transfer, where the target embodiment is absent from all training data, from pretrain-exposed zero-shot transfer, where it appears only during pretraining. We then introduce a controlled benchmark spanning 14 held-out target embodiments across simulation and real-world validation. Within this framework, we conduct a controlled analysis of four factors: state-action representations, pretraining embodiment diversity, auxiliary co-training objectives, and target-embodiment exposure. Experimental results show that local end-effector (EEF) state-action representations, the source embodiment diversity, and auxiliary co-training improve cross-embodiment transfer by around 15, 18, and 7 percentage points, respectively. We further find that adding only 5% target-embodiment data during pretraining improves average target-embodiment progress by 13.4 percentage points, showing that strict and pretrain-exposed zero-shot transfer are distinct and should be reported separately. Together, these findings provide practical guidance for evaluating and improving cross-embodiment VLA transfer in stationary tabletop manipulation with two-finger grippers, while motivating future investigation of broader settings including mobile-base control, dexterous hands, and long-horizon tasks.
发表机构
- Galbot(智元机器人)
- CFCS, School of CS, Peking University(北京大学计算机学院CFCS)
- Peking University(北京大学)
- Renmin University of China(中国人民大学)
- Xiamen University Malaysia(马来西亚厦门大学)
- The University of Hong Kong(香港大学)
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。