arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02546cs.RO

ZETA:面向桌面操作的零样本跨 embodiment VLA 迁移的受控研究

ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation

Mi Yan, Wenhao Zhang, Zhiqi Zhang, Yu Peng, Tangxinyu Wang, Lingfei Zhai, Jiayi Su, Shengliang Deng, Lin Peng, Yaowei Liu, Yuxing Chen, Zhiyuan Wei, Jilong Wang… 展开作者

Mi Yan, Wenhao Zhang, Zhiqi Zhang, Yu Peng, Tangxinyu Wang, Lingfei Zhai, Jiayi Su, Shengliang Deng, Lin Peng, Yaowei Liu, Yuxing Chen, Zhiyuan Wei, Jilong Wang, Jiayi Chen, Jiangran Lyu, Zhizheng Zhang, He Wang

首次发表
浏览论文内容

中文总结 AI 辅助

ZETA研究区分两类零样本跨embodiment VLA迁移,引入含14个目标embodiment的受控基准,发现状态-动作表示等因素可提升迁移性能,为桌面操作场景提供实用指导。

中文摘要 AI 辅助

对未见 embodiment 的零样本泛化对于可泛化的视觉-语言-动作(VLA)模型至关重要,因为机器人硬件不断发展,且特定任务的数据采集成本高昂。然而,对该问题的系统理解仍然有限,部分原因是文献缺乏统一的零样本迁移定义,以及能将 embodiment 变化与任务、场景或协议差异隔离开的受控评估设置。为解决这一差距,我们首先区分严格零样本迁移(目标 embodiment 未出现在所有训练数据中)与预训练暴露型零样本迁移(目标 embodiment 仅在预训练期间出现)。随后,我们引入了一个受控基准,涵盖模拟和真实世界验证中的14个保留目标 embodiment。在该框架内,我们对四个因素进行了受控分析:状态-动作表示、预训练 embodiment 多样性、辅助联合训练目标以及目标 embodiment 暴露。实验结果表明,局部末端执行器(EEF)状态-动作表示、源 embodiment 多样性以及辅助联合训练分别将跨 embodiment 迁移性能提升了约15、18和7个百分点。我们进一步发现,在预训练期间仅添加5%的目标 embodiment 数据,可使平均目标 embodiment 进展提升13.4个百分点,这表明严格零样本迁移与预训练暴露型零样本迁移是不同的,应分别报告。这些发现共同为评估和改进两指夹爪桌面操作场景中的跨 embodiment VLA 迁移提供了实用指导,同时推动未来对更广泛设置的研究,包括移动基座控制、灵巧手和长 horizon 任务。

英文摘要

Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) models as robot hardware evolves and task-specific data collection remains costly. However, a systematic understanding of this problem remains limited, in part because the literature lacks a unified zero-shot transfer definition and controlled evaluation settings that isolate embodiment changes from differences in tasks, scenes, or protocols. To address this gap, we first distinguish strict zero-shot transfer, where the target embodiment is absent from all training data, from pretrain-exposed zero-shot transfer, where it appears only during pretraining. We then introduce a controlled benchmark spanning 14 held-out target embodiments across simulation and real-world validation. Within this framework, we conduct a controlled analysis of four factors: state-action representations, pretraining embodiment diversity, auxiliary co-training objectives, and target-embodiment exposure. Experimental results show that local end-effector (EEF) state-action representations, the source embodiment diversity, and auxiliary co-training improve cross-embodiment transfer by around 15, 18, and 7 percentage points, respectively. We further find that adding only 5% target-embodiment data during pretraining improves average target-embodiment progress by 13.4 percentage points, showing that strict and pretrain-exposed zero-shot transfer are distinct and should be reported separately. Together, these findings provide practical guidance for evaluating and improving cross-embodiment VLA transfer in stationary tabletop manipulation with two-finger grippers, while motivating future investigation of broader settings including mobile-base control, dexterous hands, and long-horizon tasks.

发表机构

  • Galbot(智元机器人)
  • CFCS, School of CS, Peking University(北京大学计算机学院CFCS)
  • Peking University(北京大学)
  • Renmin University of China(中国人民大学)
  • Xiamen University Malaysia(马来西亚厦门大学)
  • The University of Hong Kong(香港大学)
  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑