arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.22836cs.RO

Cloak: 通过遮蔽末端执行器实现零样本跨本体操作

Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA

Michael Piseno, Guy Tevet, C. Karen Liu

首次发表
浏览论文内容

中文总结 AI 辅助

提出Cloak训练方法,通过遮蔽腕部相机中的末端执行器,使VLA模型实现零样本跨本体迁移,无需新本体数据。

中文摘要 AI 辅助

我们提出Cloak,一种训练方法,通过遮蔽末端执行器在自身腕部相机中的图像,赋予视觉-语言-动作(VLA)模型零样本跨本体迁移能力。末端执行器占据腕部视图的大片一致区域,遮蔽它允许进行与本体无关的视觉推理。Cloak根据机器人的已知几何形状,在仿真中实时精确地生成掩码,无需分割或生成模型。在训练期间,我们增强掩码,使模型能够泛化到训练时未见过的本体。我们通过Cloak-VLA演示了该方法,这是一个在单个平行夹爪数据集上使用Cloak训练的VLA模型。从未收集新本体的数据。Cloak-VLA零样本迁移到各种未见过的本体,包括另一个夹爪、另一只手臂和一只五指手,同时保持源本体的性能。通过将腕部视图与其自身本体解耦,Cloak允许数据比收集它的硬件更持久。

英文摘要

We present Cloak, a training recipe that endows a Vision-Language-Action (VLA) model with zero-shot cross-embodiment transfer by cloaking the end-effector from its own wrist camera. The end-effector occupies a large and consistent region of the wrist view and masking it allows for embodiment-agnostic visual reasoning. Cloak renders a mask in simulation from the robot's known geometry, accurately and in real time, with no segmentation or generative models. During training, we augment the mask so the model generalizes to embodiments unseen at training time. We demonstrate the recipe with Cloak-VLA, a VLA trained with Cloak on a single parallel-jaw gripper dataset. No data of new embodiments is ever collected. Cloak-VLA transfers zero-shot to various unseen embodiments, including another gripper, another arm, and a five-fingered hand, while preserving the source embodiment's performance. By decoupling the wrist view from its own embodiment, Cloak allows data to outlive the hardware it was collected on.

发表机构

  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

↑